Intelligent contract vulnerability detection method and system based on multi-modal data

By integrating multimodal data and using efficient algorithms, the problems of low accuracy, poor efficiency and insufficient interpretability in smart contract vulnerability detection are solved, and more efficient and accurate vulnerability detection and lower training costs are achieved.

CN120068081APending Publication Date: 2025-05-30TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057221.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing smart contract vulnerability detection technology has problems such as low accuracy, poor efficiency, high false alarm rate and lack of interpretability, making it difficult to effectively detect and repair potential security defects in smart contracts.

Method used

Automated smart contract vulnerability detection using a more efficient algorithm by integrating and analyzing multimodal data, including smart contract source code, vulnerability location and description information. Specific steps include obtaining raw data, preprocessing data, using large language models to generate multimodal data, training vulnerability detection models and performing detection.

Benefits of technology

It improves the accuracy and efficiency of vulnerability detection, reduces training costs, enhances the model's adaptability to complex contracts, and improves the generalization ability and interpretability of classifiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068081A_ABST
    Figure CN120068081A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent contract vulnerability detection method and system based on multi-modal data. The method comprises the following steps: acquiring and preprocessing original data; based on the preprocessed data, generating multi-modal data by using LLM; based on the multi-modal data, performing vulnerability detection model training; and performing vulnerability detection by using the trained vulnerability detection model. According to the method, the big language model is introduced, detailed description information, including vulnerability positioning, code line numbers and the like, of vulnerabilities can be automatically generated in combination with priori knowledge, the requirement for manual intervention is remarkably reduced, and the detection efficiency is improved. Besides, the output of the LLM is optimized through an intelligent prompt mechanism, so that the model can adaptively generate accurate description information in different vulnerability scenes, and the vulnerability detection precision is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of blockchain smart contract vulnerability detection, and specifically relates to a smart contract vulnerability detection method and system based on multimodal data. Background Art

[0002] With the rapid development of blockchain technology, more and more enterprises have started to use smart contracts to create complex decentralized applications, and they have been widely used in many fields such as finance, insurance, and law. Smart contract vulnerability detection is a crucial security practice in the blockchain, aiming to identify and fix potential security defects embedded in smart contracts. Smart contracts are essentially computer programs that automatically execute, control, or record contract terms, and they run on the blockchain platform, so they have the characteristics of being immutable and highly transparent. However, this also means that once a smart contract is deployed, its code and functions usually cannot be changed, and even if there are vulnerabilities, they cannot be replaced in time. Therefore, any security vulnerability or programming error may lead to serious consequences, including the loss of funds, the leakage of user privacy, or malicious attacks. On the one hand, the emergence of smart contracts has also reduced the labor cost and transaction time, and further improved the transparency and security of the blockchain. But on the other hand, due to the characteristics of smart contracts, once deployed, they cannot be modified, and any code vulnerability may cause serious consequences. In practical applications, smart contract vulnerabilities have led to many major security incidents. Any contract vulnerability or programming error may lead to serious consequences, including the loss of funds, the leakage of user privacy, or malicious attacks.

[0003] However, the current vulnerability detection technology has the following defects:

[0004] In terms of the effect of vulnerability detection, although the existing manual detection has high accuracy, the detection process is time-consuming and laborious, and the efficiency is low. At the same time, the existing machine learning-based methods show the disadvantages of low precision and high false positive rate in vulnerability detection, and they often seem powerless in the face of complex smart contract logic and are difficult to handle diverse vulnerability scenarios.

[0005] From the perspective of the detection algorithm analysis, the current vulnerability detection technology mainly relies on the analysis of the source code of vulnerabilities, and uses neural networks to learn the features in the code for detection. However, this method has significant limitations. Neural networks may not be able to effectively capture the deep features of the code, resulting in poor learning effects, and consuming a large amount of computing resources during the training process, reducing the practicability and efficiency of the algorithm.

[0006] In terms of interpretability, most of the current technologies for detecting smart contract vulnerabilities using neural networks usually lack sufficient interpretability. These models trained based on source code are mostly regarded as "black boxes", making it difficult to clearly explain the decision-making process of the models, and thus unable to provide strong guidance for subsequent experiments and improvements. Summary of the Invention

[0007] The present invention integrates and analyzes multi-modal data, such as information on smart contract source code, vulnerability location, vulnerability description, etc., and uses a more efficient algorithm to perform automated smart contract vulnerability detection, providing a smart contract vulnerability detection method and system based on multi-modal data.

[0008] To achieve the above object, the technical solutions provided by the present invention are as follows:

[0009] First aspect

[0010] The present invention provides a smart contract vulnerability detection method based on multi-modal data, including the following steps:

[0011] Step S1: Obtain the original data and perform preprocessing;

[0012] Step S2: Based on the preprocessed data, use the LLM to generate multi-modal data;

[0013] Step S3: Based on the multi-modal data, perform vulnerability detection model training;

[0014] Step S4: Use the trained vulnerability detection model to perform vulnerability detection.

[0015] Further, the step S1 specifically includes the following steps:

[0016] Step S11: Obtain the blockchain original data from the network;

[0017] Step S12: Continuously monitor the generation of new blocks, synchronize new data in real time, and perform data verification at the same time. The obtained data is cached locally and then subjected to hash verification;

[0018] Step S13: Parse the downloaded original data, extract the contract code information, decompile the bytecode of the smart contract, and generate readable source code;

[0019] Step S14: Clean up redundant or invalid data.

[0020] Further, the step S2 specifically includes the following steps:

[0021] Step S21: Construct a smart contract vulnerability detection meta buffer;

[0022] Step S22: Combine the source code data generated in Step S1 with Step S21 to customize the prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format_requirement}, which respectively represent prior knowledge, whether it is vulnerable code, vulnerability type, source code, and output format; where format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description;

[0023] Step S23: Invoke the LLM API, input the prompt defined in Step S22, and guide the large language model to output more accurate vulnerability description information;

[0024] Step S24: If the description information generated in Step S23 matches the actual code vulnerability, save the result;

[0025] Step S25: If the description generated in Step S23 does not match the actual situation, further adjust the prompt content for secondary prompt correction. If the correction is ineffective, manually label the vulnerability.

[0026] Further, the said Step S3 specifically includes the following steps:

[0027] Step S31: Integrate the data generated in Steps S24 and S25, and insert the relevant description information into the corresponding position of the source code according to the vulnerability location to form multimodal data as the training data of the vulnerability detection model;

[0028] Step S32: Use the UniXcoder model as a benchmark to initialize the vulnerability detection model;

[0029] Step S33: Fine-tune through the LoRA algorithm, and the formula is: W = W 0 +ΔW;

[0030] Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings, and these embedding vectors can effectively express the vulnerability characteristics in the code;

[0031] Step S35: Use the KAN network as a detector, and replace the traditional fixed activation function with its learnable activation function;

[0032] Step S36: Give priority to using the high-quality description information generated by one-time prompts for training. At the same time, add the description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompt strategies;

[0033] Step S37: Generate code embeddings and perform vulnerability detection After completing model fine-tuning, use the fine-tuned UniXcoder model to generate code embeddings and combine with the KAN network for final vulnerability detection.

[0034] Further, the step S4 specifically includes the following steps:

[0035] Step S41: Input the new data preprocessed in S1 into the model trained in S3 for vulnerability detection;

[0036] Step S42: Evaluate the accuracy rate and related evaluation parameters of vulnerability detection on the source code without adding any additional information.

[0037] Second aspect

[0038] Corresponding to the method, the present invention provides an intelligent contract vulnerability detection system based on multimodal data, including an original data acquisition unit, a multimodal data generation unit, a model training unit, and a vulnerability detection unit;

[0039] The original data acquisition unit is used to acquire original data and perform preprocessing;

[0040] The multimodal data generation unit is used to generate multimodal data using the LLM based on the preprocessed data;

[0041] The model training unit is used to train a vulnerability detection model based on the multimodal data;

[0042] The vulnerability detection unit is used to perform vulnerability detection using the trained vulnerability detection model.

[0043] Among them, the original data acquisition unit is specifically used to perform the following steps:

[0044] Step S11: Obtain blockchain original data from the network;

[0045] Step S12: Continuously monitor the generation of new blocks, synchronize new data in real time, and perform data verification at the same time. The obtained data is cached locally and then subjected to hash verification;

[0046] Step S13: Parse the downloaded original data, extract contract code information, decompile the bytecode of the smart contract, and generate readable source code;

[0047] Step S14: Clean up redundant or invalid data.

[0048] Among them, the multimodal data generation unit is specifically used to perform the following steps:

[0049] Step S21: Construct an intelligent contract vulnerability detection meta-buffer;

[0050] Step S22: Combine the generated source code data with Step S21, and customize the prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format_requirement}, where the distribution represents prior knowledge, whether it is vulnerable code, vulnerability type, source code, and output format; among which format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description;

[0051] Step S23: Invoke the LLM API, input the prompt defined in Step S22, and guide the large language model to output more accurate vulnerability description information;

[0052] Step S24: If the description information generated in Step S23 matches the actual code vulnerability, save the result;

[0053] Step S25: If the description produced in Step S23 does not match the actual situation, further adjust the prompt content for secondary prompt correction. If the correction is ineffective, manually annotate the vulnerability.

[0054] Among them, the model training unit is specifically used to perform the following steps:

[0055] Step S31: Integrate the data generated in Steps S24 and S25, and insert the relevant description information into the corresponding position of the source code according to the vulnerability location to form multimodal data as the training data for the vulnerability detection model;

[0056] Step S32: Initialize the vulnerability detection model with the UniXcoder model as the benchmark;

[0057] Step S33: Fine-tune through the LoRA algorithm, and the formula is: W = W 0 +ΔW;

[0058] Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings, and these embedding vectors can effectively express the vulnerability characteristics in the code;

[0059] Step S35: Use the KAN network as the detector, and replace the traditional fixed activation function with its learnable activation function;

[0060] Step S36: Preferentially use the high-quality description information generated by one-time prompting for training. At the same time, add the description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompting strategies;

[0061] Step S37: Generate code embeddings and perform vulnerability detection After completing the model fine-tuning, use the fine-tuned UniXcoder model to generate code embeddings and combine with the KAN network for final vulnerability detection.

[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0063] 1. Improve the accuracy and efficiency of vulnerability detection. In the prior art, although manual vulnerability detection has high precision, it is time-consuming and laborious; while the method based on traditional machine learning has problems of low detection accuracy and high false positive rate. By introducing a large language model, the present invention can automatically generate detailed description information of vulnerabilities, including vulnerability location, code line numbers, etc., significantly reducing the need for manual intervention and improving the detection efficiency. In addition, the output of the LLM is optimized through an intelligent prompting mechanism, enabling the model to adaptively generate accurate description information in different vulnerability scenarios, further improving the accuracy of vulnerability detection.

[0064] 2. Reduce the training cost. The present invention utilizes the LoRA low-rank fine-tuning algorithm, which significantly reduces the consumption of training resources while optimizing the model's detection ability. Compared with traditional model fine-tuning methods, LoRA not only reduces the parameters and computational overhead required during training, but also only needs to adjust 10%-20% of the original parameters to achieve an ideal optimization effect. This improvement significantly reduces the training time and the consumption of hardware resources, thereby reducing the overall training cost.

[0065] 3. Enhance the adaptability of the model to complex contracts. Existing intelligent contract vulnerability detection models often show insufficient adaptability when facing complex contract logic. By combining multi-modal data (source code and vulnerability description information) and using the UniXcoder model for fine-tuning, the present invention improves the model's understanding ability of complex intelligent contracts. Enabling the model to effectively handle tasks at different levels, thereby improving the detection ability for different types of vulnerabilities and ensuring more stable detection effects.

[0066] 4. Improve the generalization ability and detection effect of the classifier. Traditional multi-layer perceptrons (MLPs) have poor generalization ability in intelligent contract vulnerability detection and are prone to overfitting. To address this issue, the present invention eliminates the dependence on the linear weight matrix fundamentally by introducing the KAN network and using a learnable activation function to replace the fixed activation function, enabling the model to better adapt to changing input data when dealing with different types of vulnerabilities. Therefore, the present invention can maintain a high detection accuracy in different types of contract vulnerability scenarios, significantly improving the generalization ability and reliability of the classifier.

[0067] 5. Achieve higher interpretability. Most of the vulnerability detection models in the prior art are regarded as "black boxes" and lack interpretability, making it difficult to provide guidance for subsequent vulnerability repair and experiments. However, by combining the vulnerability description information generated by the LLM, the present invention enables each vulnerability detection process to produce a detailed explanation for review, helping developers understand the problems in the code more clearly and providing an effective basis for further optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 Schematic diagram of the steps of an intelligent contract vulnerability detection method based on multi-modal data provided by an embodiment of the present invention;

[0069] Figure 2 Flowchart of the operation of an intelligent contract vulnerability detection method based on multi-modal data provided by an embodiment of the present invention;

[0070] Figure 3 Flowchart of the specific implementation details of an intelligent contract vulnerability detection method based on multi-modal data provided by an embodiment of the present invention:

[0071] Figure 4 Flowchart of the LLM prompt output in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0073] As Figures 1 - 4 shown, an embodiment of the present invention provides an intelligent contract vulnerability detection method based on multi-modal data, including the following steps:

[0074] Step S1: Obtain the original data and perform preprocessing;

[0075] Step S2: Based on the preprocessed data, use the LLM to generate multimodal data;

[0076] Step S3: Based on the multimodal data, conduct vulnerability detection model training;

[0077] Step S4: Use the trained vulnerability detection model to perform vulnerability detection.

[0078] Preferably, step S1 specifically includes the following steps:

[0079] Step S11: Obtain the original blockchain data from the network. The system establishes a connection with the nodes of the network to ensure access to the real-time data in the blockchain. Select to use the official client to obtain the data.

[0080] Step S12: Continuously monitor the generation of new blocks and synchronize the new data in real time to ensure the integrity and timeliness of the data. At the same time, conduct data verification. The obtained data is cached locally and then subjected to hash verification to ensure consistency with the original data in the blockchain network and avoid data corruption or tampering during the transmission process.

[0081] Step S13: Parse the downloaded original data, extract the contract code (Bytecode) information, and decompile the bytecode of the smart contract to generate readable source code.

[0082] Step S14: Clean up redundant or invalid data, such as empty transactions, contract addresses without code, or failed transaction records, to improve the data quality.

[0083] Preferably, step S2 specifically includes the following steps:

[0084] Step S21: Construct a smart contract vulnerability detection meta buffer (a lightweight library) that contains simple descriptions of various common vulnerabilities and possible vulnerability code styles, serving as a high-level thinking template and prior knowledge for using the LLM for vulnerability detection.

[0085] Step S22: Combine the source code data generated in step S1 with step S21 to customize the prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format_requirement}, which respectively represent prior knowledge, whether it is vulnerability code, vulnerability type, source code, and output format. Among them, format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description.

[0086] Step S23: Invoke the LLM API, input the prompt defined in Step S22, and guide the large language model to output more accurate vulnerability description information.

[0087] Step S24: If the description information generated in Step S23 matches the actual code vulnerability, save the result.

[0088] Step S25: If the description generated in Step S23 does not match the actual situation, further adjust the prompt content for secondary prompt correction. If the correction is ineffective, manually label the vulnerability to ensure the effectiveness and accuracy of the generated information.

[0089] Preferably, Step S3 specifically includes the following steps:

[0090] Step S31: Integrate the data generated in Steps S24 and S25, and insert the relevant description information into the corresponding position of the source code according to the vulnerability location to form multimodal data as the training data for the vulnerability detection model.

[0091] Step S32: Use the UniXcoder model as a benchmark to initialize the vulnerability detection model.

[0092] Step S33: Fine-tune through the LoRA algorithm. The formula is: W = W 0 +ΔW. The LoRA algorithm reduces the training amount of model parameters through the low-rank matrix method, thereby reducing the computational overhead. During the fine-tuning process, the fine-tuning parameters are only 10%-20% of the original model, significantly reducing the training time and resource consumption, while improving the performance of the model in intelligent contract vulnerability detection.

[0093] Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings. These embedding vectors can effectively express the vulnerability features in the code.

[0094] Step S35: Use the KAN network as a detector, and replace the traditional fixed activation function with its learnable activation function to enhance the expression ability and generalization ability of the model.

[0095] The loss function is:

[0096]

[0097] where, N: the number of samples; C: the number of classes; y i,c This is the distribution of the true label of sample i on class c, and p^ i,c is the predicted probability that the model thinks sample i belongs to class c, usually calculated through the softmax function.

[0098] Step S36: Preferentially use the high-quality description information generated by a single prompt for training to ensure the stability of the model; meanwhile, add the description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompt strategies.

[0099] Step S37: Generate code embeddings and perform vulnerability detection After the model fine-tuning is completed, the present invention uses the fine-tuned UniXcoder model to generate code embeddings and combines with the KAN network for final vulnerability detection.

[0100] Preferably, step S4 specifically includes the following steps:

[0101] Step S41: Input the new data processed in S1 into the model trained in S3 for vulnerability detection.

[0102] Step S42: Evaluate the accuracy rate of vulnerability detection and related evaluation parameters on the source code without adding any additional information.

[0103] To ensure the accuracy of the detection results, the test data includes the description information generated by multiple prompts to verify the detection effect of the model under different prompt strategies. In contrast, the training data preferentially uses the high-quality description information that can be obtained with a single prompt to enhance the detection stability of the model. Through this step, the present invention combines code embeddings and the KAN network, successfully solves problems such as overfitting and insufficient generalization ability that occur in the traditional MLP model during the vulnerability detection process, and further improves the detection accuracy.

[0104] Corresponding to the method, the present invention provides an intelligent contract vulnerability detection system based on multimodal data, including an original data acquisition unit, a multimodal data generation unit, a model training unit, and a vulnerability detection unit;

[0105] The original data acquisition unit is used to acquire original data and perform preprocessing;

[0106] The multimodal data generation unit is used to generate multimodal data using the LLM based on the preprocessed data;

[0107] The model training unit is used to perform vulnerability detection model training based on the multimodal data;

[0108] The vulnerability detection unit is used to perform vulnerability detection using the trained vulnerability detection model.

[0109] Among them, the original data acquisition unit is specifically used to perform the following steps:

[0110] Step S11: Obtain blockchain original data from the network;

[0111] Step S12: Continuously monitor the generation of new blocks, synchronize new data in real time, and perform data verification simultaneously. The obtained data is cached locally and then subjected to hash verification;

[0112] Step S13: Parse the downloaded original data, extract contract code information, decompile the bytecode of the smart contract, and generate readable source code;

[0113] Step S14: Clean up redundant or invalid data.

[0114] Among them, the multi-modal data generation unit is specifically used to execute the following steps:

[0115] Step S21: Construct a meta buffer for detecting smart contract vulnerabilities;

[0116] Step S22: Combine the generated source code data with Step S21, and customize the prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format_requirement}, which respectively represent prior knowledge, whether it is vulnerable code, vulnerability type, source code, and output format; where format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description;

[0117] Step S23: Invoke the LLM API, input the prompt defined in Step S22, and guide the large language model to output more accurate vulnerability description information;

[0118] Step S24: If the description information generated in Step S23 matches the actual code vulnerability, save the result;

[0119] Step S25: If the description generated in Step S23 does not match the actual situation, further adjust the prompt content for secondary prompt correction. If the correction is ineffective, manually label the vulnerability.

[0120] Among them, the model training unit is specifically used to execute the following steps:

[0121] Step S31: Integrate the data generated in Steps S24 and S25, and insert relevant description information into the corresponding position of the source code according to the vulnerability location to form multi-modal data as the training data for the vulnerability detection model;

[0122] Step S32: Initialize the vulnerability detection model with the UniXcoder model as the benchmark;

[0123] Step S33: Fine-tune using the LoRA algorithm, with the formula: W = W 0 + ΔW;

[0124] Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings, and these embedding vectors can effectively represent the vulnerability features in the code;

[0125] Step S35: Use the KAN network as a detector and replace the traditional fixed activation function with its learnable activation function;

[0126] Step S36: Prioritize training using the high-quality description information generated by a single prompt. Meanwhile, add the description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompt strategies;

[0127] Step S37: Generate code embeddings and perform vulnerability detection After completing the model fine-tuning, use the fine-tuned UniXcoder model to generate code embeddings and combine with the KAN network for the final vulnerability detection.

[0128] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A smart contract vulnerability detection method based on multimodal data, characterized in that: The steps include: Step S1: obtaining raw data and preprocessing; Step S2: Generate multimodal data using LLM based on the preprocessed data; Step S3: training a vulnerability detection model based on the multimodal data; Step S4: Use the trained vulnerability detection model to perform vulnerability detection.

2. According to claim 1, a smart contract vulnerability detection method based on multimodal data is characterized in that: The step S1 specifically includes the following steps: Step S11: Obtaining blockchain raw data from the network; Step S12: Continuously monitor the generation of new blocks, synchronize new data in real time, and perform data verification at the same time. The acquired data is cached locally and then hash-verified; Step S13: Parse the downloaded raw data, extract the contract code information, decompile the bytecode of the smart contract, and generate readable source code; Step S14: Clean up redundant or invalid data.

3. According to claim 2, a smart contract vulnerability detection method based on multimodal data is characterized in that: The step S2 specifically includes the following steps: Step S21: construct a smart contract vulnerability detection meta buffer; Step S22: Combine the source code data generated in step S1 with step S21, and customize prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format at_requirement}, where the distribution represents prior knowledge, whether it is a vulnerability code, vulnerability type, source code, and output format; wherein format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description; Step S23: call the LLM API, input the prompt defined in step S22, and guide the large language model to output more accurate vulnerability description information; Step S24: If the description information generated in step S23 is consistent with the actual code vulnerability, the result is saved; Step S25: If the description produced in step S23 is inconsistent with the actual situation, further adjust the prompt content and perform a secondary prompt correction. If the correction is ineffective, manually mark the vulnerability.

4. According to claim 3, a smart contract vulnerability detection method based on multimodal data is characterized in that: The step S3 specifically includes the following steps: Step S31: Integrate the data generated in steps S24 and S25, and insert relevant description information into the corresponding position of the source code according to the vulnerability location to form multimodal data as training data for the vulnerability detection model; Step S32: using the UniXcoder model as a benchmark to initialize the vulnerability detection model; Step S33: fine-tune using the LoRA algorithm, the formula is: W = W0 + ΔW; Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings. These embedding vectors can effectively express the vulnerability characteristics in the code. Step S35: Using the KAN network as a detector, and using its learnable activation function to replace the traditional fixed activation function; Step S36: Prioritize the use of high-quality description information generated by one prompt for training, and at the same time, add description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompt strategies; Step S37: Generate code embedding and perform vulnerability detection After completing model fine-tuning, use the fine-tuned UniXcoder model to generate code embedding, and combine it with the KAN network for final vulnerability detection.

5. According to claim 4, a smart contract vulnerability detection method based on multimodal data is characterized in that: The step S4 specifically comprises the following steps: Step S41: input the new data pre-processed by S1 into the model after training in S3 for vulnerability detection; Step S42: Evaluate the accuracy of vulnerability detection and related evaluation parameters on the source code without adding any additional information.

6. A smart contract vulnerability detection system based on multimodal data, characterized in that: It includes a raw data acquisition unit, a multimodal data generation unit, a model training unit, and a vulnerability detection unit; The raw data acquisition unit is used to acquire raw data and perform preprocessing; The multimodal data generating unit is used to generate multimodal data using LLM based on the preprocessed data; The model training unit is used to perform vulnerability detection model training based on the multimodal data; The vulnerability detection unit is used to perform vulnerability detection using the trained vulnerability detection model.

7. According to claim 6, a smart contract vulnerability detection system based on multimodal data is characterized in that: The original data acquisition unit is specifically used to perform the following steps: Step S11: Obtaining blockchain raw data from the network; Step S12: Continuously monitor the generation of new blocks, synchronize new data in real time, and perform data verification at the same time. The acquired data is cached locally and then hash-verified; Step S13: Parse the downloaded raw data, extract the contract code information, decompile the bytecode of the smart contract, and generate readable source code; Step S14: Clean up redundant or invalid data.

8. According to claim 7, a smart contract vulnerability detection system based on multimodal data is characterized in that: The multimodal data generating unit is specifically used to perform the following steps: Step S21: construct a smart contract vulnerability detection meta buffer; Step S22: Combine the generated source code data with step S21, and customize prompt content P = {prior_knowledge, is_a_vulnerability, vulnerability_type, code, format_requi rement}, where the distribution represents prior knowledge, whether it is a vulnerability code, vulnerability type, source code, and output format; wherein format_requirement includes is_a_vulnerability, vulnerability_location, and vulnerability_description; Step S23: call the LLM API, input the prompt defined in step S22, and guide the large language model to output more accurate vulnerability description information; Step S24: If the description information generated in step S23 is consistent with the actual code vulnerability, the result is saved; Step S25: If the description produced in step S23 is inconsistent with the actual situation, further adjust the prompt content and perform a secondary prompt correction. If the correction is ineffective, manually mark the vulnerability.

9. The smart contract vulnerability detection system based on multimodal data according to claim 8 is characterized in that: The model training unit is specifically used to perform the following steps: Step S31: Integrate the data generated in steps S24 and S25, and insert relevant description information into the corresponding position of the source code according to the vulnerability location to form multimodal data as training data for the vulnerability detection model; Step S32: using the UniXcoder model as a benchmark to initialize the vulnerability detection model; Step S33: fine-tune using the LoRA algorithm, the formula is: W = W0 + ΔW; Step S34: Input the source code into the fine-tuned UniXcoder model to generate code embeddings. These embedding vectors can effectively express the vulnerability characteristics in the code. Step S35: Using the KAN network as a detector, and using its learnable activation function to replace the traditional fixed activation function; Step S36: Prioritize the use of high-quality description information generated by one prompt for training, and at the same time, add description information generated by multiple prompts to the test data to verify the detection effect of the model under different prompt strategies; Step S37: Generate code embedding and perform vulnerability detection After completing model fine-tuning, use the fine-tuned UniXcoder model to generate code embedding, and combine it with the KAN network for final vulnerability detection.