Fuzzy testing method and device for block chain virtual machine

By designing a prompt word iterative update mechanism in the fuzz testing of blockchain virtual machines and using a generative big model to generate targeted smart contract code, the problem of difficulty in discovering high hidden security vulnerabilities in virtual machines in the existing technology is solved, and the effectiveness of fuzz testing is improved.

CN119961933APending Publication Date: 2025-05-09SHANGHAI FENGBAO INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997328.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing blockchain virtual machine fuzz testing methods are difficult to generate effective and targeted smart contract code, making it difficult to discover highly concealed security vulnerabilities in virtual machines.

Method used

By designing a prompt word iterative update mechanism, a generative big model is used to generate diverse and targeted smart contract codes as virtual machine test input, improving the efficiency of fuzz testing.

Benefits of technology

Improves the effectiveness of blockchain virtual machine fuzz testing and increases the possibility of discovering vulnerabilities hidden in deep-level virtual machines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961933A_ABST
    Figure CN119961933A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a block chain virtual machine fuzz test method, which guides a generative large model to generate diverse and targeted smart contract codes as virtual machine test input by designing a cue word iterative update mechanism, improves the fuzz test efficiency, and improves the fuzz test accuracy. And the possibility of finding vulnerabilities hidden in the deep region of the block chain virtual machine is increased. The block chain virtual machine fuzz testing device disclosed by the embodiment of the specification also has the above beneficial effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of blockchain technology, and in particular to a blockchain virtual machine fuzzy testing method and device. Background Art

[0002] The current blockchain virtual machine security detection method mainly focuses on fuzz testing technology, which inputs a large number of smart contracts into the virtual machine and runs the virtual machine to discover its potential security vulnerabilities and abnormal behaviors. Due to the high complexity of the smart contract language and its large input space, conventional fuzz testing methods often find it difficult to generate effective and targeted smart contract codes, making it difficult to discover some highly hidden security vulnerabilities in the virtual machine. Summary of the invention

[0003] One or more embodiments of the present specification provide a blockchain virtual machine fuzz testing method and device, which can update prompt words based on a prompt word iterative update mechanism to guide the generative large model to generate more targeted test inputs, thereby improving the effectiveness of blockchain virtual machine fuzz testing.

[0004] In a first aspect, a blockchain virtual machine fuzz testing method is provided, comprising:

[0005] Based on the smart contract writing requirements, determine at least one candidate prompt word;

[0006] For each candidate prompt word, using a generative large model to generate an initial test case based on the candidate prompt word;

[0007] Pre-testing the target virtual machine using the initial test case, and determining a better prompt word from the at least one candidate prompt word according to the result of the pre-test;

[0008] At least one round of fuzz testing is performed on the target virtual machine; in each round of fuzz testing, the target virtual machine is fuzz tested by generating test cases based on the preferred prompt words using the generative large model, and the preferred prompt words are updated according to the fuzz test results according to a preset prompt word update strategy.

[0009] As an optional implementation of the method of the first aspect, based on the smart contract writing requirements, determining at least one candidate prompt word specifically includes:

[0010] Based on the smart contract writing requirements, construct smart contract writing specifications, smart contract business requirements background materials and smart contract samples;

[0011] The smart contract writing specification, the smart contract business requirement background material and the smart contract example are input into a pre-trained large language model for information extraction to obtain the at least one candidate prompt word.

[0012] As an optional implementation manner of the method of the first aspect, pre-testing the target virtual machine using the initial test case, and determining a better prompt word from the at least one candidate prompt word according to the result of the pre-test, specifically includes:

[0013] According to the fuzzy test result of the initial test case, the validity level of the candidate prompt word corresponding to the initial test case is determined by using a preset scoring strategy;

[0014] From the at least one candidate prompt word, a candidate prompt word with the highest effectiveness level is selected as the preferred prompt word.

[0015] As an optional implementation of the method of the first aspect, updating the preferred prompt word according to the fuzzy test result according to a preset prompt word update strategy specifically includes:

[0016] Based on the new abnormal result triggered in the fuzzy test result, the better prompt word is adjusted so that the updated better prompt word points to a direction that can effectively trigger the abnormal result.

[0017] Specifically, based on the abnormal result triggered in the fuzzy test result, the preferred prompt word is adjusted, specifically including:

[0018] At least one of keyword adjustment and semantic enhancement is used to adjust the preferred prompt word.

[0019] As an optional implementation manner of the method of the first aspect, the method further includes:

[0020] After generating a test case based on the preferred prompt word using the generative large model, mutating the test case to obtain a mutated test case;

[0021] The target virtual machine is fuzz tested using the mutation test case.

[0022] Specifically, updating the preferred prompt word according to the fuzzy test result according to the preset prompt word update strategy specifically includes:

[0023] If the fuzzy test result of the mutation test case triggers a new abnormal result, the bias of the better prompt word is adjusted so that the updated better prompt word points to the mutation direction corresponding to the mutation test case.

[0024] As an optional implementation manner of the method of the first aspect, the method further includes:

[0025] After generating a test case based on the preferred prompt word using the generative big model, generating a variant test case with the same semantics as the test case using the generative big model;

[0026] The target virtual machine is fuzz tested using the test case and the variant test case.

[0027] Specifically, updating the preferred prompt word according to the fuzzy test result according to the preset prompt word update strategy specifically includes:

[0028] If the fuzzy test result of the variant test case is inconsistent with the fuzzy test result of the test case, the bias of the better prompt word is adjusted so that the updated better prompt word points to the test path corresponding to the variant test case.

[0029] As an optional implementation manner of the method of the first aspect, the method further includes:

[0030] Get a pre-trained large language model;

[0031] Integrating a bidirectional adapter in the target layer of the large language model to perform bidirectional extraction and feature fusion on the output features of the target layer, and inputting the fused features into the next layer of the target layer;

[0032] Obtaining training samples for the generation task applicable to the test case;

[0033] The training samples are used to fine-tune the parameters of the bidirectional adapter to obtain the generative large model.

[0034] In the second aspect, a blockchain virtual machine fuzz testing device is provided, comprising:

[0035] A data acquisition module, configured to obtain smart contract writing requirements;

[0036] A candidate prompt word generation module, configured to determine at least one candidate prompt word based on the smart contract writing requirements;

[0037] A pre-test module is configured to generate an initial test case based on each candidate prompt word using a generative large model; pre-test a target virtual machine using the initial test case, and determine a better prompt word from the at least one candidate prompt word according to a result of the pre-test;

[0038] A fuzzy testing module is configured to perform at least one round of fuzzy testing on the target virtual machine. In each round of fuzzy testing, the generative large model is used to generate test cases based on the preferred prompt words to perform fuzzy testing on the target virtual machine.

[0039] The prompt word updating module is configured to update the preferred prompt word according to the fuzzy test result of the test case in each round of fuzzy testing according to a preset prompt word updating strategy.

[0040] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned blockchain virtual machine fuzz testing method.

[0041] In a fourth aspect, an electronic device is provided, including:

[0042] One or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, causes the electronic device to perform the above-mentioned blockchain virtual machine fuzz testing method.

[0043] The beneficial effect of the blockchain virtual machine fuzz testing method described in one or more embodiments of this specification is that the method guides the generative large model to generate diverse and targeted smart contract codes as virtual machine test input by designing a prompt word iterative update mechanism, thereby improving the efficiency of fuzz testing and increasing the possibility of discovering vulnerabilities hidden in deep-level areas of the blockchain virtual machine. The blockchain virtual machine fuzz testing device described in the embodiments of this specification also has the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 A flowchart of a blockchain virtual machine fuzz testing method provided for one or more embodiments of this specification.

[0046] Figure 2 A flowchart of a blockchain virtual machine fuzz testing method provided in one or more embodiments of this specification in a specific implementation scenario.

[0047] Figure 3 A schematic diagram of the structure of a fine-tuning training framework for a generative large model provided in one or more embodiments of this specification.

[0048] Figure 4 A schematic diagram of a training process of a bidirectional adapter provided in one or more embodiments of this specification.

[0049] Figure 5 This is a schematic diagram of the structure of a blockchain virtual machine fuzzy testing device provided in one or more embodiments of this specification.

[0050] Figure 6 A schematic diagram of the structure of an electronic device provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0052] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0053] Those skilled in the art will appreciate that the terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0054] Blockchain is essentially a decentralized distributed ledger technology that uses cryptographic methods to ensure data security, immutability and transparency, making it a reliable data storage system. At present, blockchain has been widely used in many fields such as financial services, supply chain management, and identity authentication. An important component of blockchain technology is the blockchain virtual machine, which is a virtualized computing environment running on the blockchain network. It is responsible for executing smart contracts and decentralized applications and ensuring that smart contracts are automatically executed according to predetermined rules. The blockchain virtual machine is the core environment for executing smart contracts, so it is crucial to ensure the security of the virtual machine.

[0055] Current blockchain virtual machine security detection methods mainly focus on fuzz testing technology, which is to input a large number of smart contracts into the virtual machine and run the virtual machine to discover its potential security vulnerabilities and abnormal behaviors. Fuzz testing technology focuses on mutating the smart contract bytecode or high-level language source code before compilation to generate a large number of different smart contract codes to test the virtual machine. However, current blockchain virtual machine fuzz testing methods still face some key problems: (1) Due to the high complexity of the smart contract language and its large input space, the fuzz testing method cannot generate effective and targeted smart contract code, making it difficult to discover some highly hidden security vulnerabilities in the virtual machine; (2) Due to ignoring the complex state changes during the contract execution process, the generated test input may miss some deep-level state-dependent vulnerabilities; (3) The test input generated by mutation during the fuzz testing process may produce code format errors, which will cause the virtual machine to crash, making it impossible to fully test the virtual machine.

[0056] In view of this, one or more embodiments of this specification propose a blockchain virtual machine fuzzy testing method and device to at least partially solve the above technical problems.

[0057] The blockchain virtual machine fuzzy testing method and device described in one or more embodiments of this specification will be further described in detail below in conjunction with the drawings and specific embodiments of the specification, but this detailed description does not constitute a limitation on the embodiments of this specification.

[0058] Please refer to Figure 1 , Figure 1 This is a flowchart of a blockchain virtual machine fuzzy testing method proposed in one or more embodiments of this specification. The blockchain virtual machine fuzzy testing method can be applied to a fuzz tester to implement fuzzy testing of the blockchain virtual machine. Figure 1 As shown, the method may include steps S100 to S106.

[0059] S100: Determine at least one candidate prompt word based on smart contract writing requirements.

[0060] S102: For each candidate prompt word, generate an initial test case based on the candidate prompt word using a generative large model.

[0061] S104: Pre-testing the target virtual machine using the initial test case, and determining a better prompt word from at least one candidate prompt word according to the result of the pre-test.

[0062] S106: Perform at least one round of fuzz testing on the target virtual machine. In each round of fuzz testing, the generative large model is used to generate test cases based on the better prompt words to perform fuzz testing on the target virtual machine, and the better prompt words are updated according to the fuzz testing results according to the preset prompt word update strategy.

[0063] The above blockchain virtual machine fuzz testing method, by designing a prompt word iterative update mechanism, guides the generative large model to generate effective smart contract code as a test case for the virtual machine, effectively guiding the fuzz tester to discover vulnerabilities hidden in the deep areas of the virtual machine. This method can generate diverse and targeted smart contract codes, improving the efficiency of blockchain virtual machine fuzz testing.

[0064] The following will be combined with specific implementation methods. Figure 1 The blockchain virtual machine fuzz testing method shown is explained.

[0065] S100: Determine at least one candidate prompt word based on smart contract writing requirements.

[0066] The above-mentioned smart contract writing requirements mainly include the logical structure of the smart contract, functional module division, coding standards, and business requirements of the business scenarios in which the smart contract is applied. By building or obtaining smart contract writing requirements, it can be ensured that the generated smart contract code meets specific business requirements and security requirements.

[0067] In this step S100, at least one candidate prompt word is determined based on the smart contract writing requirements, which is essentially to extract information from the smart contract writing requirements, so as to obtain candidate prompt words that can characterize the characteristics of the smart contract code required by them. Based on this, a pre-trained large language model can be used to extract information from the smart contract writing requirements, so as to obtain the above-mentioned at least one candidate prompt word.

[0068] Please refer to Figure 2 In some implementations, a distillation-based large model can be selected as the above-mentioned large language model, and the pre-collected smart contract writing specification samples, smart contract business demand background material samples and smart contract samples are used to train the distillation-based large model for information extraction tasks, wherein the smart contract writing specification samples are used to describe the logical structure, functional module division and coding standards of the smart contract, the smart contract business demand background material samples are used to describe the business requirements of the business scenario of the smart contract, and the smart contract samples are smart contract samples that meet the requirements of the above-mentioned smart contract writing specification samples and smart contract business demand background material samples. The trained distillation-based large model can generate the above-mentioned candidate prompt words based on the input smart contract writing specification, smart contract business demand background material and smart contract samples.

[0069] It should be noted that in step S100 , the method of generating candidate prompt words and the number of candidate prompt words can be adaptively set according to needs, and this embodiment does not limit this.

[0070] S102: For each candidate prompt word, generate an initial test case based on the candidate prompt word using a generative large model.

[0071] The above-mentioned generative large model refers to a pre-trained large language model, which can generate smart contract code as an initial test case based on the input prompt information.

[0072] In some embodiments, a pre-trained large language model can be directly obtained, and then a bidirectional adapter can be integrated into the target layer of the large language model to form the above-mentioned generative large model. The bidirectional adapter can perform bidirectional extraction and feature fusion on the output features of the target layer, and input the fused features into the next layer of the target layer to enhance the modeling ability of the large language model for the context. Finally, the above-mentioned generative large model is fine-tuned using the pre-collected training samples for the test case generation task. During the fine-tuning process, the parameters of the large language model can be frozen to avoid comprehensive back propagation of the large language model, and only the parameters of the bidirectional adapter are trained, thereby effectively reducing computational costs and resource consumption.

[0073] Please refer to Figure 3 , Figure 3 A structural diagram of a fine-tuning training framework for a generative large model is shown. Figure 3 The large language model in the text adopts the transformer structure, which mainly includes four parts: input layer, encoder, decoder and output layer. When the large language model is applied, only the one-decoder structure is used, that is, the single decoder structure. Therefore, in order to simplify the description, Figure 3 Only the input layer, decoder and output layer are shown.

[0074] like Figure 3 As shown, the bidirectional adapter can be integrated in different target layers of the decoder, for example, in the middle layer and the terminal layer of the decoder. The bidirectional adapter can perform bidirectional information integration on the input at different levels to enhance the modeling ability of the entire generative model for the context and improve the generation quality of the generative model.

[0075] The above-mentioned bidirectional adapter is a lightweight neural network, which can be functionally divided into a downstream adapter and an upstream adapter. The upstream adapter is used to encode the input features from left to right, and the downstream adapter is used to encode the input features from right to left. Structurally, the bidirectional adapter mainly includes a bottleneck layer, a cross-layer adapter, and a bidirectional encoding layer. The bottleneck layer is used to compress the input information through a low-dimensional hidden layer, and then restore it through a fully connected layer of increasing dimensions. This design can capture information while keeping the model lightweight. The cross-layer adapter is used to extract and combine bidirectional context information across multiple layers, thereby enhancing the modeling ability of the context. The bidirectional encoding layer is used to encode information from left to right and from right to left at the same time. The context vectors in the left and right directions can be spliced ​​together, or integrated through the attention mechanism to generate a comprehensive context representation.

[0076] In order to enhance the generative large model's ability to model context, different training tasks may be constructed to fine-tune the above-mentioned bidirectional adapter. These training tasks may be adaptively constructed according to demand, and this embodiment does not impose any restrictions on this.

[0077] Please refer to Figure 4 ,In some implementations, the above training tasks can be constructed as a contract sequential ,code modeling task, a mask modeling task, and a bidirectional sequence ,generation task.

[0078] The contract sequence code modeling task is to predict the context of the smart contract code by the generative big model given the context of the smart contract code. The bidirectional adapter can use the context on the left and right sides to help the generative big model make more accurate predictions.

[0079] The mask modeling task is to mask part of the code in the smart contract and input it into the generative model, which predicts the masked code based on the bidirectional context. This type of task can effectively train the bidirectional adapter to capture global context information.

[0080] The task of bidirectional sequence generation is to give a part of the content in the middle of the smart contract, and let the generative model predict the context of the smart contract based on the given content. In this task, the bidirectional adapter can improve the accuracy and logic of the generated contract code by combining bidirectional information.

[0081] After completing the fine-tuning of the bidirectional adapter parameters, you can also build more specific downstream tasks, such as classification, generation, or translation, to jointly train the bidirectional adapter and the large language model. When conducting joint training for specific downstream tasks, you can unfreeze some of the weights of the large language model, allowing it to be fine-tuned together with the bidirectional adapter to achieve better task adaptability and performance improvement. At the same time, you can also appropriately introduce regularization techniques and model pruning methods to prevent the large language model from overfitting, thereby achieving collaborative optimization of the bidirectional adapter and the large language model, and training a generative large model with better performance.

[0082] S104: Pre-testing the target virtual machine using the initial test case, and determining a better prompt word from at least one candidate prompt word according to the result of the pre-test.

[0083] In the pre-testing phase, the initial test cases can be used to perform a small-scale fuzz test on the target virtual machine to capture boundary conditions, abnormal states, and potential vulnerabilities during the execution of the smart contract, and verify the execution of the generated smart contract code in the target virtual machine.

[0084] According to the test results of the pre-test, a pre-built scoring mechanism can be used to score the candidate prompt words, and a better prompt word can be selected according to the scoring results. Specifically, according to the fuzzy test results of the initial test case, the validity level of the candidate prompt words corresponding to the initial test case can be determined using a preset scoring strategy, and then from at least one candidate prompt word, the candidate prompt word with the highest validity level can be selected as the better prompt word.

[0085] The scoring criteria of the above-mentioned scoring strategy can be adaptively set according to the needs, and this embodiment does not limit this. For example, the scoring criteria may include indicators such as the validity of generating test inputs, code coverage, and whether potential problems or vulnerabilities are found. For example, if a new abnormal result is found in a fuzzy test, the candidate prompt word is valid, and 1 point can be added to the candidate prompt word. If this abnormal result contains an uncovered code path, 1 point can be added to the candidate prompt word. If a new vulnerability or defect is also found in this abnormal result, the candidate prompt word is added 1 point. If no abnormal result is found in a fuzzy test, the candidate prompt word is not added. Finally, the candidate prompt words are sorted in order from high to low according to the score, and the candidate prompt word ranked first is taken as the better prompt word.

[0086] S106: Perform at least one round of fuzz testing on the target virtual machine. In each round of fuzz testing, the generative large model is used to generate test cases based on the better prompt words to perform fuzz testing on the target virtual machine, and the better prompt words are updated according to the fuzz testing results according to the preset prompt word update strategy.

[0087] The above-mentioned round of fuzz testing includes at least one comprehensive fuzz testing, that is, in a round of fuzz testing, only one fuzz testing can be performed, or multiple fuzz testing can be performed. In the first round of fuzz testing, the initial test cases corresponding to the better prompt words in the pre-test phase can be directly used as the test cases of the current round. It is also possible to adjust the current better prompt words according to the pre-test results of the initial test cases corresponding to the better prompt words in the pre-test phase, and then use the generative large model to generate new test cases based on the updated better prompt words. These new test cases can have different structures, logics and parameter combinations.

[0088] It should be noted that the number of rounds of fuzz testing can be set according to demand, and this embodiment does not limit this. For example, it can be set to end the above-mentioned blockchain virtual machine fuzz testing method when the fuzz testing time reaches a preset time threshold.

[0089] In each round of fuzz testing, the preferred prompt word can be updated according to the fuzz testing results according to the preset prompt word update strategy. Specifically, if in a round of fuzz testing, the fuzz testing results trigger new abnormal results, such as discovering new vulnerabilities or triggering uncovered code paths, it means that the preferred prompt word has a better test effect on the specific path. At this time, the preferred prompt word can be adjusted so that the updated preferred prompt word points to the direction that can effectively trigger the abnormal result.

[0090] The above prompt word update strategy can be set according to needs, and then the fuzzy test results are matched with the rules in the prompt word update strategy to determine the specific optimal prompt word update method. For example, at least one adjustment method of keyword adjustment and semantic enhancement can be used to adjust the optimal prompt word.

[0091] Specifically, if the current fuzz test results find new vulnerabilities, the types of vulnerabilities can be identified, such as buffer overflow, SQL injection, cross-site scripting (XSS), etc. By identifying different types of vulnerabilities, the keywords in the current better prompt words are updated and increased. For example, if a certain type of input causes a vulnerability, the relevant keywords can be included in the prompt words to increase the coverage of the better prompt words. If the current fuzz test results find uncovered paths, the semantic expression of the prompt words can be optimized to make it clearer and more precise. For example, descriptive words or phrases can be added to help the model understand the context of the vulnerability.

[0092] In some embodiments, in each round of fuzz testing, the generative large model can be directly used to generate test cases based on the current preferred prompt words, and then the test cases are used to perform fuzz testing on the target virtual machine, and according to the results of this fuzz test, the current preferred prompt words are updated according to the preset prompt word update strategy. For example, if this fuzz test finds new abnormal results, such as discovering new vulnerabilities or triggering uncovered code paths, the fine-grained control of the current preferred prompt words can be adjusted, such as adjusting the bias of input generation, or further refining the description of the prompt words to guide the generation of more targeted new inputs. If this fuzz test does not find new abnormal results, the pointing range of the current preferred prompt words can be adjusted so that the generative large model can generate test cases with a wider test range.

[0093] In some implementations, after generating a test case based on the current preferred prompt word using the generative large model, the test case can be mutated to obtain a mutated test case, and then the mutated test case can be used to perform fuzz testing on the target virtual machine. If the fuzzy test result of the mutated test case triggers a new abnormal result, such as discovering a new vulnerability or triggering an uncovered code path, it means that the current preferred prompt word has a better mutation effect on the specific path. At this time, the bias of the preferred prompt word can be adjusted so that the updated preferred prompt word points to the mutation direction corresponding to the mutated test case, thereby further optimizing the mutation strategy and enabling the generative large model to generate more targeted mutated test cases.

[0094] The specific methods for mutating the test cases mentioned above may include modifying input parameters, adjusting contract logic, or changing the order of function calls. Specifically, new smart contract instructions can be inserted into the code of an existing smart contract to increase the overall complexity of the contract. It is also possible to modify the existing code portion of the smart contract to increase the complexity of the code while keeping the control flow as stable as possible. Modification operations may include changing variables, operators, or instruction sequences, but it is necessary to ensure that these changes do not disrupt the basic logic and flow of the program.

[0095] After generating mutation test cases, you can also perform semantic checking to verify the correctness of the mutated test case code. This process ensures that the code is valid after mutation and correctly reflects the expected functionality. Semantic checking involves analyzing the mutated code to confirm that there are no syntax errors, the logical structure is preserved, and the program behavior has not been inappropriately changed.

[0096] In some embodiments, after using the generative big model to generate a test case based on the current preferred prompt word, the generative big model can be used to generate a variant test case with the same semantics as the test case, such as by reorganizing statements, replacing synonyms, or adjusting the order of parameters, to generate test cases of different forms but logically equivalent, and then use the test case and the variant test case to perform fuzz testing on the target virtual machine.

[0097] Analyze the execution results of the generated semantically identical test cases and variant test cases in the target virtual machine, and focus on checking whether the variant test cases lead to the same behavior as the test cases. If inconsistent execution results occur, the current preferred prompt word needs to be optimized. For example, the bias of the preferred prompt word can be adjusted so that the updated preferred prompt word points to the test path corresponding to the variant test case. This update enables the generative large model to more effectively generate diverse but semantically equivalent inputs, further covering different test paths.

[0098] The above is Figure 1 A blockchain virtual machine fuzz testing method is shown in the figure. The method proposes a prompt word screening method, which selects the better prompt words for fuzz testing by prioritizing different prompt words and the smart contract codes generated by them, thereby improving the efficiency of fuzz testing. The method proposes a joint fine-tuning method of a bidirectional adapter and a large speech model, which constructs a generative large model by building and training a bidirectional adapter and inserting it into different layers of a large language model for joint fine-tuning, thereby enhancing the bidirectional information capture capability of the generative large model in the generation task and improving the effectiveness of smart contract code generation. The method proposes a prompt word update strategy, which continuously optimizes the prompt words during the fuzz testing process to help the generative large model generate more effective and targeted smart contract codes, thereby improving the test input quality of the blockchain virtual machine. The method proposes a blockchain virtual machine fuzz testing scheme guided by a generative large model, which can generate effective virtual machine test inputs and increase the possibility of discovering vulnerabilities hidden in deep areas of the virtual machine.

[0099] Corresponding to the above-mentioned blockchain virtual machine fuzzy testing method, one or more embodiments of this specification also propose a blockchain virtual machine fuzzy testing device, which is suitable for a fuzz tester. Figure 5 , Figure 5 This is a schematic diagram of the structure of a blockchain virtual machine fuzzy testing device proposed in one or more embodiments of this specification. The device can be used to implement the above-mentioned blockchain virtual machine fuzzy testing method. It should be noted that the blockchain virtual machine fuzzy testing method described in one or more embodiments of this application can rely on Figure 5 The blockchain virtual machine fuzz testing device shown is implemented, but not limited to this device.

[0100] like Figure 5 As shown, the blockchain virtual machine fuzzy testing device includes:

[0101] The data acquisition module 501 is configured to obtain smart contract writing requirements.

[0102] The candidate prompt word generation module 502 is configured to determine at least one candidate prompt word based on the smart contract writing requirements.

[0103] The pre-test module 503 is configured to generate an initial test case based on each candidate prompt word using a generative large model; pre-test the target virtual machine using the initial test case, and determine a better prompt word from at least one candidate prompt word according to the result of the pre-test.

[0104] The fuzz testing module 504 is configured to perform at least one round of fuzz testing on the target virtual machine. In each round of fuzz testing, the generative large model is used to generate test cases based on the preferred prompt words to perform fuzz testing on the target virtual machine.

[0105] The prompt word updating module 505 is configured to update the better prompt words according to the fuzzy test results of the test cases in each round of fuzzy testing according to a preset prompt word updating strategy.

[0106] For the above-mentioned data acquisition module 501, the smart contract writing requirements it obtains mainly include the logical structure of the smart contract, the functional module division, the coding standard, and the business requirements of the business scenario in which the smart contract is applied. By constructing or obtaining the smart contract writing requirements, it can ensure that the generated smart contract code meets specific business requirements and security requirements.

[0107] Specifically, the above-mentioned smart contract writing requirements may include smart contract writing specifications, smart contract business demand background materials and smart contract samples. Among them, the smart contract writing specifications are used to describe the logical structure, functional module division and coding standards of smart contracts, the smart contract business demand background materials are used to describe the business requirements of the business scenarios of smart contracts, and the smart contract samples are smart contracts that meet the requirements of the above-mentioned smart contract writing specifications and smart contract business demand background materials.

[0108] For the candidate prompt word generation module 502, in some embodiments, the candidate prompt word generation module 502 can use a pre-trained large language model to extract information from the smart contract writing requirements, so as to obtain the above-mentioned at least one candidate prompt word. For example, the candidate prompt word generation module 502 can select a distillation-type large model as the above-mentioned large language model, and then input the obtained smart contract writing specifications, smart contract business requirements background materials and smart contract examples into the distillation-type large model to obtain the above-mentioned at least one candidate prompt word.

[0109] Regarding the above-mentioned pre-test module 503, during the pre-test phase, the module can use the initial test cases to perform a small-scale fuzz test on the target virtual machine to capture the boundary conditions, abnormal states and potential vulnerabilities during the execution of the smart contract, and verify the execution of the generated smart contract code in the target virtual machine.

[0110] According to the test results of the pre-test, the pre-test module 503 can use a pre-built scoring mechanism to score the candidate prompt words, and select a better prompt word according to the scoring result. Specifically, according to the fuzzy test results of the initial test case, the validity level of the candidate prompt words corresponding to the initial test case can be determined using a preset scoring strategy, and then from at least one candidate prompt word, the candidate prompt word with the highest validity level is selected as the better prompt word.

[0111] The scoring criteria of the above-mentioned scoring strategy can be adaptively set according to the needs, and this embodiment does not limit this. For example, the scoring criteria may include indicators such as the validity of generating test inputs, code coverage, and whether potential problems or vulnerabilities are found. For example, if a new abnormal result is found in a fuzzy test, the candidate prompt word is valid, and 1 point can be added to the candidate prompt word. If this abnormal result contains an uncovered code path, 1 point can be added to the candidate prompt word. If a new vulnerability or defect is also found in this abnormal result, the candidate prompt word is added 1 point. If no abnormal result is found in a fuzzy test, the candidate prompt word is not added. Finally, the candidate prompt words are sorted in order from high to low according to the score, and the candidate prompt word ranked first is taken as the better prompt word.

[0112] For the above-mentioned fuzzy test module 504, each round of fuzzy testing of the fuzzy test module 504 includes at least one comprehensive fuzzy test, that is, in one round of fuzzy testing, the fuzzy test module 504 can perform only one fuzzy test or multiple fuzzy tests. In the first round of fuzzy testing, the initial test cases corresponding to the better prompt words in the pre-test phase can be directly used as the test cases of the current round. It is also possible to adjust the current better prompt words according to the pre-test results of the initial test cases corresponding to the better prompt words in the pre-test phase, and then use the generative large model to generate new test cases based on the updated better prompt words. These new test cases can have different structures, logics and parameter combinations.

[0113] In each round of fuzz testing, the prompt word update module 505 can update the better prompt word according to the fuzz testing result of the fuzz testing module 504 and the pre-deployed prompt word update strategy. Specifically, if in a round of fuzz testing, the fuzz testing result of the fuzz testing module 504 triggers a new abnormal result, such as discovering a new vulnerability or triggering an uncovered code path, it means that the better prompt word has a better test effect on the specific path. At this time, the prompt word update module 505 can adjust the better prompt word so that the updated better prompt word points to a direction that can effectively trigger the abnormal result.

[0114] The above prompt word update strategy can be set according to needs, and then the fuzzy test results are matched with the rules in the prompt word update strategy to determine the specific optimal prompt word update method. For example, at least one adjustment method of keyword adjustment and semantic enhancement can be used to adjust the optimal prompt word.

[0115] Specifically, if the current fuzzy test results find new vulnerabilities, the prompt word update module 505 can identify the type of vulnerability, such as buffer overflow, SQL injection, cross-site scripting (XSS), etc. By identifying different types of vulnerabilities, the keywords in the current better prompt words are updated and added. For example, if a certain type of input causes a vulnerability, the prompt word update module 505 can include the relevant keywords in the prompt words to improve the coverage of the better prompt words. If the current fuzzy test results find uncovered paths, the prompt word update module 505 can optimize the semantic expression of the prompt words to make them clearer and more accurate, for example, adding descriptive words or phrases to help the model understand the context of the vulnerability.

[0116] In some embodiments, in each round of fuzz testing, the fuzz testing module 504 can directly use the generative large model to generate test cases based on the current better prompt words, and then use the test cases to perform fuzz testing on the target virtual machine. The prompt word update module 505 updates the current better prompt words according to the preset prompt word update strategy based on the results of this fuzz test. For example, if this fuzz test finds new abnormal results, such as finding new vulnerabilities or triggering uncovered code paths, the prompt word update module 505 can adjust the fine-grained control of the current better prompt words, such as adjusting the bias of input generation, or further refining the description of the prompt words to guide the generation of more targeted new inputs. If this fuzz test does not find new abnormal results, the prompt word update module 505 can adjust the pointing range of the current better prompt words so that the generative large model can generate test cases with a wider test range.

[0117] In some embodiments, the fuzzy test module 504 may also mutate the test case after generating the test case based on the current preferred prompt word using the generative large model to obtain a mutated test case, and then use the mutated test case to perform fuzzy testing on the target virtual machine. If the fuzzy test result of the mutated test case triggers a new abnormal result, such as discovering a new vulnerability or triggering an uncovered code path, it means that the current preferred prompt word has a better mutation effect on the specific path. At this time, the prompt word update module 505 can adjust the bias of the preferred prompt word so that the updated preferred prompt word points to the mutation direction corresponding to the mutated test case, thereby further optimizing the mutation strategy and enabling the generative large model to generate more targeted mutated test cases.

[0118] The specific methods by which the fuzz testing module 504 mutates the test cases may include modifying input parameters, adjusting contract logic, or changing the order of function calls. Specifically, the fuzz testing module 504 may insert new smart contract instructions into the code of an existing smart contract, aiming to increase the overall complexity of the contract. The fuzz testing module 504 may also modify the existing code portion of the smart contract to increase the complexity of the code while keeping the control flow as stable as possible. The modification operations performed by the fuzz testing module 504 may include changing variables, operators, or instruction sequences, but it is necessary to ensure that these changes do not destroy the basic logic and flow of the program.

[0119] After generating the mutated test case, the fuzz testing module 504 can also perform semantic checking to verify the correctness of the mutated test case code. This process ensures the validity of the mutated code and correctly reflects the expected functionality. Semantic checking involves analyzing the mutated code to confirm that there are no syntactic errors, the logical structure is preserved, and the program behavior has not been improperly changed.

[0120] In some embodiments, after using the generative big model to generate a test case based on the current preferred prompt word, the fuzz testing module 504 can also use the generative big model to generate a variant test case with the same semantics as the test case, for example, by reorganizing statements, replacing synonyms, or adjusting the order of parameters, to generate test cases of different forms but logically equivalent, and then use the test case and the variant test case to perform fuzzy testing on the target virtual machine.

[0121] The prompt word update module 505 analyzes the execution results of the generated semantically identical test cases and variant test cases in the target virtual machine, and focuses on checking whether the variant test case leads to the same behavior as the test case. If inconsistent execution results occur, the current better prompt word needs to be optimized. For example, the bias of the better prompt word can be adjusted so that the updated better prompt word points to the test path corresponding to the variant test case. This update enables the generative large model to more effectively generate diverse but semantically equivalent inputs, further covering different test paths.

[0122] For the above-mentioned blockchain virtual machine fuzz testing device, taking a module as an example of a software functional unit, the data acquisition module 501 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the data acquisition module 501 may include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with similar geographical locations. Among them, usually a region may include multiple AZs.

[0123] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0124] As an example of a hardware functional unit, the data acquisition module 501 may include at least one computing device, such as a server, etc. Alternatively, the data acquisition module 501 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0125] The multiple computing devices included in the data acquisition module 501 can be distributed in the same region or in different regions. The multiple computing devices included in the data acquisition module 501 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the data acquisition module 401 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0126] In other embodiments, the data acquisition module 501 can be used to execute any step in the above-mentioned blockchain virtual machine fuzzy testing method, the candidate prompt word generation module 502 can be used to execute any step in the above-mentioned blockchain virtual machine fuzzy testing method, the pre-test module 503 can be used to execute any step in the above-mentioned blockchain virtual machine fuzzy testing method, and the prompt word update module 505 can be used to execute any step in the above-mentioned blockchain virtual machine fuzzy testing method. The steps that the data acquisition module 501, the candidate prompt word generation module 502, the pre-test module 503, the fuzzy test module 504 and the prompt word update module 505 are responsible for implementing can be specified as needed, and the data acquisition module 501, the candidate prompt word generation module 502, the pre-test module 503, the fuzzy test module 504 and the prompt word update module 505 respectively implement different steps in the above-mentioned blockchain virtual machine fuzzy testing method to realize all the functions of the above-mentioned blockchain virtual machine fuzzy testing device.

[0127] In this implementation, the blockchain virtual machine fuzzy testing device can also be applied to computing devices such as computers and servers, or to a computing device cluster including at least one computing device, to realize the specific functions of the blockchain virtual machine fuzzy testing device.

[0128] In some embodiments, an electronic device is also provided. Figure 6 , the electronic device includes: a bus 601, a processor 602, a memory 603 and a communication interface 604. The processor 602, the memory 603 and the communication interface 604 communicate with each other through the bus 601. The electronic device can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the electronic device.

[0129] The bus 601 may be a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The use of only one line does not mean that there is only one bus or one type of bus. Bus 601 may include a path for transmitting information between various components of the electronic device (eg, processor 602, memory 603, and communication interface 604).

[0130] The processor 602 may include any one or more of a processor such as a CPU, a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0131] The memory 603 may include a volatile memory, such as a random access memory (RAM). The memory 603 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0132] The memory 603 stores executable program code, and the processor 602 executes the executable program code to implement the functions of the aforementioned blockchain virtual machine fuzzy testing device, that is, to implement the aforementioned blockchain virtual machine fuzzy testing method.

[0133] The communication interface 604 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the electronic device and other devices or a communication network.

[0134] In some embodiments, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned blockchain virtual machine fuzz testing method.

[0135] The computer-readable storage medium may be any available medium that can be stored by an electronic device or a data storage device such as a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the electronic device to execute the above-mentioned blockchain virtual machine fuzz testing method.

[0136] It is to be understood that the structure illustrated in the embodiments of this specification does not constitute a specific limitation on the system of the embodiments of this specification. In other embodiments of the specification, the above system may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0137] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0138] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0139] It should be noted that the above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.

Claims

1. A blockchain virtual machine fuzz testing method, comprising: Based on the smart contract writing requirements, determine at least one candidate prompt word; For each candidate prompt word, using a generative large model to generate an initial test case based on the candidate prompt word; Pre-testing the target virtual machine using the initial test case, and determining a better prompt word from the at least one candidate prompt word according to the result of the pre-test; At least one round of fuzz testing is performed on the target virtual machine. In each round of fuzz testing, the generative large model is used to generate test cases based on the preferred prompt words to perform fuzz testing on the target virtual machine, and the preferred prompt words are updated according to the fuzz testing results according to a preset prompt word update strategy.

2. According to the method of claim 1, based on the smart contract writing requirements, determining at least one candidate prompt word specifically comprises: Based on the smart contract writing requirements, construct smart contract writing specifications, smart contract business requirements background materials and smart contract samples; The smart contract writing specification, the smart contract business requirement background material and the smart contract example are input into a pre-trained large language model for information extraction to obtain the at least one candidate prompt word.

3. The method according to claim 1, using the initial test case to pre-test the target virtual machine, and determining a better prompt word from the at least one candidate prompt word according to the result of the pre-test, specifically comprising: According to the fuzzy test result of the initial test case, the validity level of the candidate prompt word corresponding to the initial test case is determined by using a preset scoring strategy; From the at least one candidate prompt word, a candidate prompt word with the highest effectiveness level is selected as the preferred prompt word.

4. The method according to claim 1, updating the preferred prompt word according to the fuzzy test result according to a preset prompt word update strategy, specifically comprising: Based on the new abnormal result triggered in the fuzzy test result, the better prompt word is adjusted so that the updated better prompt word points to a direction that can effectively trigger the abnormal result.

5. The method according to claim 4, adjusting the preferred prompt word based on the abnormal result triggered in the fuzzy test result, specifically comprising: At least one of keyword adjustment and semantic enhancement is used to adjust the preferred prompt word.

6. The method according to claim 1, further comprising: After generating a test case based on the preferred prompt word using the generative large model, mutating the test case to obtain a mutated test case; The target virtual machine is fuzz tested using the mutation test case.

7. The method according to claim 6, updating the preferred prompt word according to the fuzzy test result according to a preset prompt word update strategy, specifically comprising: If the fuzzy test result of the mutation test case triggers a new abnormal result, the bias of the better prompt word is adjusted so that the updated better prompt word points to the mutation direction corresponding to the mutation test case.

8. The method according to claim 1, further comprising: After generating a test case based on the preferred prompt word using the generative big model, generating a variant test case with the same semantics as the test case using the generative big model; The target virtual machine is subjected to fuzz testing using the test case and the variant test case.

9. The method according to claim 8, updating the preferred prompt word according to the fuzzy test result according to a preset prompt word update strategy, specifically comprising: If the fuzzy test result of the variant test case is inconsistent with the fuzzy test result of the test case, the bias of the better prompt word is adjusted so that the updated better prompt word points to the test path corresponding to the variant test case.

10. The method according to claim 1, further comprising: Get a pre-trained large language model; Integrating a bidirectional adapter in the target layer of the large language model to perform bidirectional extraction and feature fusion on the output features of the target layer, and inputting the fused features into the next layer of the target layer; Obtaining training samples for the generation task applicable to the test case; The training samples are used to fine-tune the parameters of the bidirectional adapter to obtain the generative large model.

11. A blockchain virtual machine fuzz testing device, comprising: A data acquisition module, configured to obtain smart contract writing requirements; A candidate prompt word generation module, configured to determine at least one candidate prompt word based on the smart contract writing requirements; A pre-test module is configured to generate an initial test case based on each candidate prompt word using a generative large model; Pre-testing the target virtual machine using the initial test case, and determining a better prompt word from the at least one candidate prompt word according to the result of the pre-test; A fuzzy testing module is configured to perform at least one round of fuzzy testing on the target virtual machine. In each round of fuzzy testing, the generative large model is used to generate test cases based on the preferred prompt words to perform fuzzy testing on the target virtual machine. The prompt word updating module is configured to update the preferred prompt word according to the fuzzy test result of the test case in each round of fuzzy testing according to a preset prompt word updating strategy.

12. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 10.

13. An electronic device comprising: one or more processors; And a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the electronic device executes the method according to any one of claims 1 to 10.