Fuzzy testing method, device and equipment for distributed database

By designing test case mutation rules and test predictions for distributed databases, the problem of insufficient defect detection in distributed database testing by existing tools is solved, and efficient distributed database defect detection is achieved.

CN120832307APending Publication Date: 2025-10-24TIANJIN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511118949.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing fuzz testing tools are difficult to effectively discover defects in distributed databases related to distributed mechanisms such as data sharding, operator pushdown, and data replication. In addition, the testing cost is high and the detection logic defects are insufficient.

Method used

By using the test case generation function of SQLancer, we can generate test cases for distributed databases through non-equivalent mutation and equivalent mutation, add statements related to the distribution mechanism, trigger the distributed features, and detect the correctness of the output results through test prediction.

Benefits of technology

It achieves efficient discovery of defects related to distributed mechanisms in distributed databases, overcomes the shortcomings of existing tools, and improves testing efficiency and detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832307A_ABST
    Figure CN120832307A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed databases, in particular to a distributed database fuzzy testing method, device and equipment, and the distributed database fuzzy testing method comprises the following steps: generating an original seed by using a test case generation function of SQLancer or a test case set of a to-be-tested database; performing non-equivalent variation on the original seeds, using cases capable of triggering a distributed mechanism of the database to be tested, and performing equivalent variation to generate a pair of functional equivalent cases; and sending the varied seeds into the database to be tested, executing the varied seeds, obtaining an output result, and detecting the correctness of the output result through a test oracle. According to the method, the defect of close association with the related mechanism of the distributed database can be explored, and the defect that an existing fuzzy test tool lacks a test function for the related mechanism of the distributed database is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed database, and particularly relates to a distributed database fuzzing method, device and equipment. BACKGROUND

[0002] With the development of the Internet, the data that modern software needs to process and store gradually increases. For the consideration of efficiency and disaster recovery, distributed databases have become important underlying components of data-intensive applications. However, the architecture of the distributed database is more complex than that of the traditional database, and some defects and vulnerabilities may be introduced in the implementation process. These defects and vulnerabilities may cause great loss to the user's business, so it is necessary to fully test the implementation of the distributed database. The huge internal state space of the database makes it difficult to comprehensively test it, so in recent years, fuzzing has become a testing method that has attracted much attention. Fuzzing randomly generates a large number of test sample inputs to the system under test, and then detects the correctness of the output of the system under test through test oracle to check whether there are defects.

[0003] Compared with the traditional database, the distributed database introduces mechanisms specific to distributed systems, such as data sharding, operator pushdown, and data replication. It also introduces defects related to distributed mechanisms, so it is necessary to propose a fuzzing framework for distributed databases to explore these defects.

[0004] SQLancer is a currently existing framework for fuzzing traditional databases. The framework first generates test cases that conform to the syntax of the database under test according to manually written generation rules, and then sends them to the database under test for execution, and checks the correctness of the execution results through test oracle. If the execution result is correct, new test cases will be generated for testing, otherwise the test case that triggers the defect and the error information will be stored in the log. SQLancer can be used with test oracle to detect logical defects in the database. For example, PQS randomly selects a row of data from the query result, then synthesizes a query statement that satisfies the row of data and executes it. If the query result does not contain the selected data row, an error is reported. NoREC proposes a metamorphic relationship for the optimization of the database, which converts a query statement that may trigger optimization into a query statement that is semantically equivalent and does not trigger optimization, so that the mutated statement is equivalent to the original statement and can trigger the database optimization mechanism. TLP uses the ternary logic in the database to construct a metamorphic relationship. A query statement is split into three query statements that contain logical judgment conditions with true values of 1, 0 and NULL according to a certain attribute. The three query statements together are semantically equivalent to the original query statement.

[0005] Jepsen is a chaos testing framework for distributed systems. The framework injects defects into the system under test, such as node crashes and network partitions, and then uses the transaction consistency detector Elle and the linear consistency detector Porcupine to detect whether the system under test meets the consistency constraints it declares. Elle mainly focuses on the isolation constraint of distributed transactions, and finds defects by capturing the dependency relationship between operations and constructing a dependency graph; Porcupine is a state transition model written by the user, and then detects whether the operation history of the system under test conforms to the model to verify whether the system meets the linear consistency.

[0006] SQLancer, PQS, NoREC and TLP only design use case generation rules and test predictions for traditional databases, so their support for distributed databases is weak. In the design of test case generation rules and test predictions, they do not consider the data sharding, operator pushdown, and data replication mechanisms of distributed databases, so it is difficult to find defects related to the mechanisms of distributed systems.

[0007] Jepsen requires defect injection, and the cost of learning and using is high; in addition, the detectors Elle and Porcupine used by Jepsen require that various operations of the system under test be converted into key-value operations, and have high requirements for the structure of test cases. When testing distributed databases with these tools, not only a lot of effort is required, but also the detection of logical defects such as inconsistency of database query results will be lacking because they mainly focus on the consistency of distributed systems. SUMMARY

[0008] The present application provides a distributed database fuzzing test method, device and equipment to solve the problems in the background art.

[0009] In a first aspect, the present application provides a distributed database fuzzing test method, comprising: generating an original seed using the test case generation function of SQLancer or the test case set of the database under test; performing non-equivalent mutation on the original seed, using the test case to trigger the distributed mechanism of the database under test, and then performing equivalent mutation to generate a pair of functionally equivalent test cases; sending the mutated seed to the database under test for execution and obtaining the output result, and detecting the correctness of the output result by test prediction.

[0010] Further, after generating the original seed using the test case generation function of SQLancer or the test case set of the database under test, the method further comprises: When the original seed is from the test case set of the to-be-tested database, the seed is parsed by using the corresponding syntax parser of the to-be-tested database, and the variable name in the seed is rewritten as a variable name conforming to the context semantics.

[0011] Further, the non-equivalent variation of the original seed includes: A statement related to the distributed mechanism is added to the original seed, so as to trigger the distributed characteristics of the to-be-tested database.

[0012] Further, the adding of the statement related to the distributed mechanism to the original seed includes: The partition mode of a data table in the original seed is changed, and data is inserted into the data table according to the partition information, so as to ensure that each partition of each data table has data; According to the partition information of the partition table, a range query condition for the partition column is added to the query condition, so that the query process involves multiple data partitions.

[0013] Further, the rules of the equivalent variation include: For the same data table, the full set of shards obtained by using different partition modes is the same; For a statement, the result set obtained by disabling the optimization mechanism related to the distributed characteristics is the same as the original result set; For a statement, if the data table queried by the statement has multiple copies, the result set obtained by querying any two copies is the same.

[0014] Further, the detection of the correctness of the output result by the test oracle includes: The test oracle compares whether the execution results of the pair of functionally equivalent test cases are consistent, to determine whether the execution results are correct.

[0015] Further, after the detection of the correctness of the output result by the test oracle, the method further includes: The seed triggering the defect is returned to the seed pool, so as to update the seed pool.

[0016] In a second aspect, the present application provides a distributed database fuzzing device, which includes: A seed generation module is configured to generate an original seed by using the test case generation function of SQLancer or the test case set of the to-be-tested database. A seed variation module is configured to perform non-equivalent variation on the original seed, so that the test case can trigger the distributed mechanism of the to-be-tested database, and then perform equivalent variation to generate a pair of functionally equivalent test cases. The execution test module is configured to send the mutated seed to the to-be-tested database for execution, obtain an output result, and detect the correctness of the output result by using a test oracle.

[0017] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the distributed database fuzz testing method as described above when executing the computer program.

[0018] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the distributed database fuzz testing method as described above.

[0019] The above technical solutions of the present application have the following advantages: The distributed database fuzz testing method provided by the first aspect of the present application generates original seeds by using the test case generation function of SQLancer or the test case set of the to-be-tested database, performs non-equivalent mutation on the original seeds, uses the test cases to trigger the distributed mechanism of the to-be-tested database, performs equivalent mutation to generate a pair of functionally equivalent test cases, sends the mutated seeds to the to-be-tested database for execution and obtains an output result, detects the correctness of the output result by using a test oracle, designs test case mutation rules and test oracles specific to the distributed database mechanism, explores defects closely related to the distributed database mechanism, and overcomes the deficiency of the existing fuzz testing tools lacking test functions specific to the distributed database mechanism.

[0020] It can be understood that the beneficial effects of the above-mentioned second aspect, third aspect and fourth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the description of the embodiments or the prior art. Obviously, the drawings described below are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0022] Figure 1 The flow chart of the distributed database fuzz testing method provided by the present application; Figure 2 The principle diagram of the distributed database fuzz testing method provided by the present application; Figure 3 The example diagram of the mutation rule provided by the present application; Figure 4A structural diagram of a distributed database fuzz testing device provided in the present application is shown in FIG. 1. Figure 5 A structural diagram of an electronic device provided in the present application is shown in FIG. 1. DETAILED DESCRIPTION

[0023] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular architectures, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0024] It is to be understood that the terminology "includes", "has", "holds", "contains" or "comprising", "including" when used in this specification and in the following claims, specifies the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0025] In addition, in the description of the present application and in the following claims, the terms "first", "second", "third", etc. are used only for distinguishing between similar objects, and cannot be understood as indicating or implying relative importance.

[0026] In the present specification, the reference to "one embodiment" or "some embodiments" etc. means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of the phrases "in one embodiment", "in some embodiments", "in other embodiments", "in additional embodiments", etc. in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise specifically noted. The terms "comprising", "including", "containing", and "having" together with their conjugates mean "including but not limited to", unless otherwise expressly specified. The term "consisting of" means "including and limited to". The term "consisting essentially of" means that the composition, method, or structure contains those indicated elements as well as any additional element not materially affecting the basic and novel characteristics of the composition, method, or structure.

[0027] The purpose of the present application is to provide a distributed database fuzz testing method, which overcomes the deficiency of the existing fuzz testing tools lacking of testing functions for the mechanisms related to the distributed database, by designing test case variation rules and test prediction for the mechanisms specific to the distributed database, and exploring defects closely related to the mechanisms of the distributed database.

[0028] The specific embodiments of the present application are described in further detail below with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0029] AsFigure 1 As shown, the embodiment of the present application provides a distributed database fuzzing method, and specifically includes the following steps: generating an original seed by using a test case generation function of SQLancer or a test case set of a to-be-tested database; performing non-equivalent mutation on the original seed, using the test case to trigger a distributed mechanism of the to-be-tested database, and then performing equivalent mutation to generate a pair of functionally equivalent test cases; sending the mutated seed to the to-be-tested database for execution and obtaining an output result, and detecting the correctness of the output result by using a test oracle.

[0030] In some embodiments, after the step of generating an original seed by using a test case generation function of SQLancer or a test case set of a to-be-tested database, the method further includes: when the original seed is from the test case set of the to-be-tested database, parsing the seed by using a corresponding grammar parser of the to-be-tested database, and rewriting variable names in the seed into variable names conforming to context semantics.

[0031] In some embodiments, the step of performing non-equivalent mutation on the original seed, using the test case to trigger a distributed mechanism of the to-be-tested database, includes: adding a statement related to the distributed mechanism to the original seed, so as to trigger the distributed characteristics of the to-be-tested database.

[0032] In some embodiments, the step of adding a statement related to the distributed mechanism to the original seed includes: changing a partitioning mode of a data table in the original seed and inserting data into the data table according to partitioning information, so as to ensure that each partition of each data table contains data; and adding a range query condition for a partition column to a query condition according to partition information of a partition table, so as to make a query process involve multiple data partitions.

[0033] In some embodiments, the rules of the equivalent mutation include: for the same data table, a full set of shards obtained by using different partitioning modes is the same; for a statement, a result set obtained by disabling an optimization mechanism related to the distributed characteristics is the same as an original result set; and for a statement, if a data table queried by the statement has multiple copies, a result set obtained by querying any two copies of the data table is the same.

[0034] In some embodiments, the step of detecting the correctness of the output result by using a test oracle includes: comparing, by using the test oracle, whether execution results of the pair of functionally equivalent test cases are consistent, to determine whether the execution results are correct.

[0035] In some embodiments, after the step of detecting the correctness of the output result by using a test oracle, the method further includes: sending a seed triggering a defect back to a seed pool, to update the seed pool.

[0036] As Figure 2As shown, the distributed database fuzz testing method proposed in the present application first generates an initial seed using a public test case set and the test case generation function of SQLancer, then mutates the original seed using mutation rules, sends the mutated test case to the system under test for execution, and finally uses the test oracle to detect the correctness of the results.

[0037] The original seed is generated using the test case generation function of SQLancer or the test case set of the database under test. When the original seed comes from the test case set of the database under test, the corresponding syntax parser of the database under test is used to parse the seed, and the variable names in the seed are rewritten to conform to the context semantics.

[0038] The original seed is mutated, and the main ways of mutation are two: adding statements related to the distributed mechanism to the original seed, triggering the non-equivalent mutation of the distributed characteristics of the system under test, and through the metamorphic relationship, the equivalent mutation of the original seed is equivalent to the equivalent mutation. The present application first uses non-equivalent mutation, and the example can trigger the distributed mechanism of the system under test, and then uses equivalent mutation to generate a pair of functionally equivalent test cases, so that the test oracle can compare the result sets of the two to judge the correctness of the execution results.

[0039] The non-equivalent mutation proposed in the present application has two main ways: changing the partitioning method of the data table in the original seed and inserting data into the data table according to the partitioning information, or changing the set queried by the query statement. The former is to mutate the statement that builds the database state, and the latter is to mutate the statement that queries the data. Since the main feature of a distributed database is to divide the data into multiple copies and scatter them to multiple nodes for storage, and operations on multiple storage nodes are more likely to have defects than operations on a single storage node, the main purpose of the mutation rules designed by the present application is to scatter the data as much as possible to multiple storage nodes, and to involve as much as possible read operations on multiple storage nodes when querying.

[0040] As shown, the present application proposes two non-equivalent mutation rules: Figure 3 1. Table partition mutation. Change the partitioning method of the data table and insert data into the table according to the partitioning information, ensuring that each partition of each data table has data. Figure 3 An example of this mutation rule is given: change a non-partitioned table to a range partitioned table, and insert data into each partition according to the partition information.

[0041] 2. Query set mutation. According to the partition information of the partition table, add range query conditions for the partition column to the query condition, so that the query process involves multiple data partitions. Figure 3 ​An example of the mutation rule is given: add a query condition for column c1 to the original query, so that the query that can only find data in one partition becomes a query that finds data in multiple partitions.

[0042] After non-equivalent mutation, the present application further uses equivalent mutation on the mutated use case to generate a pair of functionally equivalent test cases. As shown in Figure 3 , the equivalent mutation rule has the following three types: 1. Partition equivalence. For the same data table, the full set of shards obtained by using different partitioning methods (i.e. and , where n1 and n2 are the number of shards) should satisfy .

[0043] 2. Optimization equivalence. For a statement q, the statement q -op obtained by disabling the optimization mechanism related to the distributed feature should satisfy -op R(q)=R(q -op ).

[0044] 3. Replication equivalence. For a statement q, if the data table queried by the statement has rep_cnt copies in total, then the result sets R(q rep1 ) and R(q rep2 ) obtained by querying any two copies should satisfy R(q rep1 )=R(q rep2 ).

[0045] The mutated seeds are sent to the database under test for execution, and the output results are obtained. The correctness of the output results of the test cases is detected by the test oracle. Whether the execution results are correct is determined by comparing whether the test cases obtained by equivalent mutation are consistent. The seeds that trigger defects are sent back to the seed pool. The seeds that have triggered defects continue to be mutated, which may trigger more defects, so sending these seeds to the seed pool can increase the probability of discovering new defects.

[0046] The main difference between a distributed database and a traditional database is that data is distributed to multiple storage nodes for separate storage. The key point of the present application is to design and implement a test case generation method with high defect triggering capability, to mutate the use cases using test case mutation rules for distributed databases, so that database operations such as update and query involve as many shards of data on storage nodes as possible, and to design and implement a test result judgment method with high capture capability, to design equivalent metamorphic relationships for mechanisms such as data sharding and data replication in distributed databases, and to implement test oracle to check the consistency of the output results of the test cases before and after metamorphic relationship mutation, to verify whether the database under test has defects.

[0047] This application can discover defects related to unique distributed database mechanisms such as table sharding, shard pruning, and operation pushdown. SQLancer focuses on traditional database testing, while Jepsen focuses on distributed system testing in a broad sense. Therefore, both are relatively weak in discovering defects related to these unique distributed database mechanisms. This application addresses these shortcomings.

[0048] Corresponding to the distributed database fuzzy testing method described in the above embodiment, Figure 4 As shown, the embodiment of the present application also provides a distributed database fuzzy testing device, which includes: The seed generation module is used to generate the original seed using the test case generation function of SQLancer or the test case set of the database to be tested; A seed mutation module is used to perform non-equivalent mutation on the original seed, so that a use case can trigger the distributed mechanism of the database under test, and then perform equivalent mutation to generate a pair of functionally equivalent use cases; The execution test module is used to send the mutated seed to the database to be tested for execution and obtain the output result, and detect the correctness of the output result through test prediction.

[0049] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0050] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0051] The present application also provides an electronic device, such as Figure 5As shown, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the distributed database fuzz testing method provided in the first aspect are implemented.

[0052] In applications, electronic devices may include, but are not limited to, processors and memories. Figure 5 These are merely examples of electronic devices and do not limit the scope of the electronic device. The electronic device may include more or fewer components than shown, or a combination of certain components or different components. For example, input / output devices, network access devices, etc. Input / output devices may include cameras, audio capture / playback devices, display screens, etc. Network access devices may include a network module for establishing wireless network connections with external devices.

[0053] In applications, a processor may be a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0054] In applications, in some embodiments, memory can be an internal storage unit of an electronic device, such as a hard drive or memory. In other embodiments, memory can also be an external storage device of the electronic device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash memory card. Memory can also include both internal storage units and external storage devices. Memory is used to store operating systems, applications, boot loaders, data, and other programs, such as computer program code. Memory can also be used to temporarily store data that has been output or is about to be output.

[0055] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0056] The present application implements all or part of the processes in the above-described method embodiments by instructing the relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When executed by a processor, the computer program may implement the steps of each of the above-described method embodiments. The computer program includes computer program code, which may be in source code form, object code form, an executable file, or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk.

[0057] Those skilled in the art will appreciate that the devices and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0058] In the embodiments provided herein, it should be understood that the disclosed devices and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interface, or the devices may be indirectly coupled or communicated in some manner, whether electrical, mechanical, or otherwise.

[0059] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method of distributed database fuzz testing, the method comprising: The method comprises the following steps: generating an original seed by using a test case generation function of SQLancer or a test case set of a database to be tested; performing non-equivalent mutation on the original seed, using an example to trigger a distributed mechanism of the database to be tested, and then performing equivalent mutation to generate a pair of functionally equivalent cases; sending the mutated seed into the database to be tested for execution and obtaining an output result, and detecting the correctness of the output result by using a test oracle.

2. The distributed database fuzz testing method of claim 1, wherein, After the step of generating an original seed by using a test case generation function of SQLancer or a test case set of a database to be tested, the method further comprises the following steps: when the original seed is from the test case set of the database to be tested, parsing the seed by using a syntax parser corresponding to the database to be tested, and rewriting variable names in the seed into variable names conforming to context semantics.

3. The distributed database fuzz testing method of claim 1, wherein, The step of performing non-equivalent mutation on the original seed, using an example to trigger a distributed mechanism of the database to be tested, comprises the following steps: adding a statement related to the distributed mechanism to the original seed, so as to trigger the distributed characteristics of the database to be tested.

4. The distributed database fuzz testing method of claim 3, wherein, The step of adding a statement related to the distributed mechanism to the original seed comprises the following steps: changing a partitioning mode of a data table in the original seed and inserting data into the data table according to partitioning information, so as to ensure that each partition of each data table contains data; according to partitioning information of a partition table, adding a range query condition for a partition column to a query condition, so as to make a query process involve multiple data partitions.

5. The distributed database fuzz testing method of claim 1, wherein, The rules of the equivalent mutation comprise the following rules: for the same data table, a full set of shards obtained by using different partitioning modes is the same; for a statement, a result set obtained by disabling an optimization mechanism related to distributed characteristics is the same as an original result set; for a statement, if a data table queried by the statement has multiple copies, a result set obtained by querying the data table on any two copies is the same.

6. The distributed database fuzz testing method of claim 1, wherein, The step of detecting the correctness of the output result by using a test oracle comprises the following step: comparing, by using a test oracle, whether execution results of the pair of functionally equivalent cases are consistent, to determine whether the execution results are correct.

7. The distributed database fuzz testing method of claim 1, wherein, After the step of detecting the correctness of the output result by using a test oracle, the method further comprises the following step: sending a seed triggering a defect back to a seed pool, to update the seed pool.

8. A distributed database fuzzing device, comprising: The method comprises the following steps: a seed generation module is configured to generate an original seed by using a test case generation function of SQLancer or a test case set of a database to be tested; a seed mutation module is configured to perform non-equivalent mutation on the original seed, using an example to trigger a distributed mechanism of the database to be tested, and then performing equivalent mutation to generate a pair of functionally equivalent cases; an execution test module is configured to send the mutated seed into the database to be tested for execution and obtain an output result, and detect the correctness of the output result by using a test oracle.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the distributed database fuzz testing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the distributed database fuzz testing method according to any one of claims 1 to 7.