System and method for fuzz testing
By designing a fuzz testing system that includes a fuzzer data generator, an input subsystem, a fuzz testing agent and a fuzzer evaluation function block, the problem of low testing speed, efficiency and quality in the prior art is solved, and the ability to efficiently detect software security vulnerabilities and errors is achieved.
Patent Information
- Application Number
- CN202380057475.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2023-08-03
- Publication Date
- 2025-05-06
AI Technical Summary
Existing fuzz testing systems cannot provide fast, efficient and high-quality testing sufficiently, making it difficult to effectively detect security vulnerabilities and errors in software.
A fuzz testing system is designed that includes a fuzzer data generator, an input subsystem, a fuzzing test agent and a fuzzer evaluation function block. The system continuously generates data units and adds hooks to the binary executable file to monitor the impact of input data on the program, evaluates and adjusts the test data in real time to improve test efficiency and quality.
Fast, efficient and high-quality fuzz testing is realized, which can effectively detect security vulnerabilities and errors in the software, and improve the efficiency and accuracy of software testing.
Smart Images

Figure CN119948465A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to the field of software testing, and in particular to systems and methods for fuzz testing (fuzzing).
[0002] background
[0003] In programming and software development, fuzzing or fuzz testing is an automated software testing technique that involves providing invalid, unexpected, or random data as input to a computer program. The program is then monitored for anomalies such as crashes, failed built-in code assertions, or potential memory leaks. Unfortunately, current fuzz testing systems do not provide sufficiently fast, efficient, and high-quality testing.
[0004] Overview
[0005] Additional features and advantages of the invention will become apparent from the following drawings and description.
[0006] In some examples, a system for fuzz testing is provided, the system comprising a fuzzer data generator configured to continuously generate data units. In some examples, the system comprises a first input subsystem, the first input subsystem configured to input each of the generated data units into a device under test, an input end of the first input subsystem communicates with an output end of the fuzzer data generator, and an output end of the first input subsystem communicates with the device under test.
[0007] In some examples, the system includes a first fuzz testing agent configured to add each of one or more hooks to a corresponding point of interest in one or more predetermined points of interest in a binary executable file that runs on a device under test, wherein in response to an input data unit, each hook outputs information associated with the corresponding point of interest, the output information including data stored in a corresponding address of a memory associated with the corresponding point of interest.
[0008] In some examples, the system includes a fuzzer evaluation functionality configured to receive information from each of the one or more hooks.
[0009] In some examples, the fuzzifier data generator is in communication with the fuzzifier evaluation functional block, and the fuzzifier data generator generates the data unit in response to an output of the fuzzifier evaluation functional block.
[0010] Unless otherwise limited, all technical and scientific terms used herein have the same meanings as those of ordinary skill in the art to which the present invention belongs. If conflict occurs, this patent application specification (including definitions) is dominant. As used herein, unless the context clearly indicates otherwise, the articles "a" and "an" refer to "at least one" or "one or more". As used herein, "and / or" refers to any one or more items in the list connected by "and / or". As an example, "x and / or y" refers to any element in the set {(x), (y), (x, y)} of three elements. In other words, "x and / or y" refers to "x, y or both x and y". As another example, "x, y and / or z" refers to any element in the set {(x), (y), (z), (x, y), (x, z), (y, z), (x, y, z)} of seven elements.
[0011] In addition, unless expressly stated to the contrary, "or" refers to an inclusive or rather than an exclusive or. For example, any of the following satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).
[0012] In addition, "a" or "an" is used to describe elements and components of embodiments of the inventive concept. This is done only for convenience and to give a general sense of the inventive concept, and "a" and "an" are intended to include one or at least one, and the singular also includes the plural unless it is obvious that it has other meanings.
[0013] As used herein, the term "approximately," when referring to a measurable value (e.g., an amount, a duration, etc.), is meant to encompass deviations of + / -10%, more preferably + / -5%, even more preferably + / -1%, and still more preferably + / -0.1% from the specified value, as such deviations are suitable for performing the disclosed apparatus and / or methods.
[0014] The following embodiments and aspects thereof of systems, tools and methods are described and illustrated, which are intended to be exemplary and illustrative, rather than limiting in scope. In various embodiments, one or more of the above problems have been reduced or eliminated, while other embodiments are directed to other advantages or improvements. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] For a better understanding of the invention and to show how it may be put into practice, reference will now be made, by way of example only, to the accompanying drawings in which like reference numerals designate corresponding parts or elements throughout.
[0017] With specific reference now to the drawings in detail, it is emphasized that the details shown are presented by way of example and only for the purpose of illustrative discussion of the preferred embodiments of the invention and in order to provide what is believed to be the most useful and easily understood description of the principles and conceptual aspects of the invention. In this regard, no attempt is made to show the structural details of the invention in more detail than is necessary for a basic understanding of the invention, and the description taken in conjunction with the drawings makes it apparent to those skilled in the art how several forms of the invention may be embodied in practice. In the drawings:
[0018] Figures 1A to 1D Various portions of an example of a fuzz testing system according to some examples of the present disclosure are shown;
[0019] FIG. 2A to FIG. 2B Shown is the Figures 1A to 1C A neural network setup of a system generating data unit;
[0020] FIG. 3A to FIG. 3B Shows the combination FIG. 2A to FIG. 2B Various high-level block diagrams of examples of fuzz testing systems for neural network devices;
[0021] FIG. 3C to FIG. 3E Shows the description Figure 3A Various diagrams of the method of operation of the fuzz testing system;
[0022] Figure 4A A high-level block diagram illustrating an example of a fuzz testing system according to some examples of the present disclosure is shown;
[0023] Figure 4B Shows Figure 4A A more detailed example of a fuzz testing system;
[0024] FIG. 4C to FIG. 4E Various high-level flow charts of fuzz testing methods utilizing network-level fuzz testing and function-level fuzz testing are shown;
[0025] Figure 4F A high-level block diagram is shown, which illustrates the placement of hooks 111 throughout the call tree;
[0026] Figure 4G A high-level block diagram of a fuzz testing agent including multiple event handlers according to some examples of the present disclosure is shown;
[0027] Figure 5A A high-level block diagram illustrating an example of a fuzz testing system according to some examples of the present disclosure is shown;
[0028] Figure 5B A high-level block diagram illustrating an example of a fuzz testing system according to some examples of the present disclosure is shown;
[0029] 6A to 6F Various high-level block diagrams showing examples of a proxy-based fuzz testing system;
[0030] Figure 7 A high-level flow chart showing a signal-based fuzz testing method according to some examples of the present disclosure; and
[0031] Figure 8 A high-level flow chart of a method of determining statistical independence of signals according to some examples of the present disclosure is shown.
[0032] Detailed description of some embodiments
[0033] In the following description, various aspects of the present disclosure will be described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the different aspects of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without the specific details presented herein. In addition, well-known features may be omitted or simplified to avoid obscuring the present disclosure. In the accompanying drawings, similar reference numerals always refer to similar parts. In order to avoid excessive confusion due to having too many reference numerals and leads on a particular figure, some components will be introduced by one or more figures without being clearly identified in each subsequent figure containing the component.
[0034] Figure 1A A high-level block diagram of a system 10 for fuzz testing is shown. In some examples, the system 10 for fuzz testing includes a fuzzer data generator 20, an input subsystem 30, a fuzz testing agent 40, and a fuzzer evaluation function block 50. In some examples, the fuzz testing agent 40 includes a timestamp generator 60. As known to those skilled in the art, the timestamp generator 60 generates a timestamp. In another example, as will be described below, the timestamp generator 60 is external to the fuzz testing agent 40.
[0035] In some examples, such as Figure 1C As shown, the security vulnerability testing system 10 includes at least one processor 70 and a memory 80. In such an example, a plurality of instructions are stored in the memory 80, and when the plurality of instructions are executed by the at least one processor 70, the at least one processor 70 performs the functions of the fuzzer data generator 20, the input subsystem 30, the fuzz test agent 40, and the fuzzer evaluation functional block 50. Therefore, in such an example, the fuzzer data generator 20, the input subsystem 30, the fuzz test agent 40, and the fuzzer evaluation functional block 50 are each composed of a corresponding instruction set stored on the memory 80.
[0036] As used herein, the terms "fuzzer data generator," "fuzz testing agent," and "fuzzer evaluation functional block" refer to various parts of a fuzzer.
[0037] In some examples, such as Figure 1D As shown, the system 10 for fuzz testing is implemented to cooperate with a test tool 100, such as the CANoe software tool commercially available from Vector Informatik GmbH of Stuttgart, Germany. In some examples, the test tool 100 includes various simulations of a network access interface and a simulated electronic control unit (ECU). In another example, the fuzzer evaluation function block 50 is implemented on the test tool 100.
[0038] The fuzz testing agent 40 adds one or more hooks 111 to the binary executable file 110, each hook 111 being added to a corresponding predetermined point of interest in the binary executable file 110. As used herein, the term "hook 111" refers to one or more lines of code that change the operation of the binary executable file 110 at the point where the hook 111 is located. In some examples, as will be described below, each hook 111 branches to the fuzz testing agent 40. In some examples, the binary executable file 110 is a binary executable file 110 of a device under test (DUT) 115 tested at the test tool 100. As used herein, the term "binary executable file" refers to a file of a machine language designed for a corresponding processor, that is, as known to those skilled in the art, a binary executable file contains executable code represented by specific processor instructions.
[0039] In some examples, the timestamp generator 60 is part of the DUT 115 or the test tool 100. In such examples, the fuzz testing agent 40 optionally communicates with the timestamp generator 60 and requests timestamps as needed from the timestamp generator 60. In another example, each hook 111 requests a timestamp from the timestamp generator 60 when activated.
[0040] In some examples, hook 111 is added by replacing an opcode at a corresponding point of interest with a branch instruction to branch to fuzz testing agent 40. In another example, hook 111 is added by overwriting an address of a corresponding point of interest in a procedure linkage table (PLT) associated with binary executable file 110. In some examples, fuzz testing agent 40 adds one or more hooks 111 to binary executable file 110 without recompiling binary executable file 110.
[0041] In some examples, the fuzz testing agent 40 is embedded within the binary executable file 110. The following describes an example of embedding the fuzz testing agent 40 with the binary executable file 110, however, this is not meant to be limiting in any way, and any known embedding method may be used without exceeding the scope of the present disclosure.
[0042] In some examples, where the source code of the binary executable file 110 is available and the binary executable file 110 is provided in an executable and linkable (ELF) format, the embedding of the fuzz test agent 40 into the binary executable file 110 is accomplished by analyzing the file by a preparation script to find available space to which the fuzz test agent 40 can fit. If there is enough space in an existing segment, a portion of the PROGBITS of the fuzz test agent 40 (i.e., a portion of the program content) is copied into a binary program image within the available space. When copying the PROGBITS of the fuzz test agent 40, the relative distances between different sections within the fuzz test agent 40 are preferably maintained. In particular, sections in an ELF file that contain various types of data and are loaded at runtime need to be mapped to addresses in CPU memory. As known to those skilled in the art in the presence of the present invention, mapping is performed by segments. Each segment contains a series of consecutive PROGBITS sections that are loaded together to the address specified by the segment. Therefore, the added segment for the fuzz testing agent 40 will load the added PROGBITS section into the process address space at runtime.
[0043] If there is not enough space in the existing segments, two new segments are added to the ELF file by the preparation script. The first segment is for read-only executable text and the second segment is for read-write access. Then, the sections of the fuzz test agent 40 are added to the added segments. Specifically, the read-write access PROGBITS section includes data and a global offset table (GOT). All segments of the ELF file are listed in the program header table. After the two new segments are added, the program header table no longer adapts to its original offset. Therefore, the program header table is moved to the end of the ELF file by the preparation script. Then the third segment is added to the program header table by the preparation script, which is arranged to load the program header table from its new position into the process address space at runtime to allow the process to be loaded and executed. The code is position-independent, so as long as the relative distance between different sections is maintained, relocation within the address space does not require any modification. However, sometimes there are global offsets in the code. These offsets are stored in the GOT and modified by the preparation script to reflect the relocation of the address.
[0044] The input subsystem 30 includes software and / or firmware input to the binary executable 110 of the DUT 115. Specifically, an input of the input subsystem 30 communicates with an output of the fuzzer data generator 20, and an output of the input subsystem 30 communicates with the DUT 115. In some examples, the input subsystem 30 includes a network interface.
[0045] In some examples, as described below, the input subsystem 30 can input data units (such as data packets) in two ways: inputting data units through a network interface for network-level fuzz testing; and inputting data units through an emulator for function-level fuzz testing. As used herein, the term "network-level fuzz testing" refers to fuzz testing of an instrumented binary executable file of a device using emulation of an ECU and / or various ports and devices as known to those skilled in the art. As used herein, the term "function-level fuzz testing" refers to: using an emulator that contains memory states associated with the process of reaching a specific function; and then providing data units directly to the function.
[0046] In some examples, the fuzzer evaluation function block 50 is not embedded in the binary executable file 110. In some examples, the fuzzer evaluation function block 50 communicates with the fuzz testing agent 40 and the fuzzer data generator 20. Although the fuzzer evaluation function block 50 and the fuzzer data generator 20 are described separately herein, this is not meant to be limited to two separate and distinct elements. In some examples, the fuzzer evaluation function block 50 and the fuzzer data generator 20 are part of a set of combined software instructions and operate as a single program.
[0047] In some examples, such as Figure 1D As shown, the fuzzer evaluation function block 50 is in communication with a network 120 suitable for cloud-based computing. In some examples, the network 120 is part of the Internet. In another example, the fuzzer evaluation function block 50 is combined with a cloud-based computing platform.
[0048] In operation, in some examples, the fuzzer evaluation function block 50 receives one or more points of interest in the binary executable file 110 from a user input (not shown). In another example, the fuzzer evaluation function block 50 scans the binary executable file 110 to identify one or more points of interest. It should be noted that these are not exclusive options, and the fuzzer evaluation function block 50 can identify points of interest in response to both: user input; and scanning of the binary executable file 110. In some examples, the fuzzer evaluation function block 50 scans the binary executable file 110 to determine a known application programming interface (API). For the Automotive Open System Architecture (AUTOSAR), this can include, for example, CanIf_RxIndication.
[0049] Other points of interest may include, but are not limited to, any one or a combination of the following: runtime environment (RTE) interface; internal application functions; AUTOSAR callouts; and various library function-like vectors, such as VStdLib_MemCpy; security-related functions, such as functions accessing a hardware security module (HSM) or libcrypto; predetermined sensitive functions, such as memcpy; parsers; conditional logic; points in the flow starting from the input entry, such as read, rxIndication, processPacket, memcpy, etc.
[0050] In some examples, the fuzzer evaluation function block 50 also defines event information that may be useful, such as: a hook 111 hit counter, i.e., how many times a specific hook 111 is reached; a notification when the value of a specific register is equal to an expected value; a notification about a corrupted memory stack; and a notification of a heap overflow. In some examples, the event information type is defined and / or approved by the user.
[0051] The above is described in an example where the fuzzer evaluation function block 50 defines points of interest and events, however, this is not meant to be limiting in any way. In another example, Figure 1B As shown, the system 10 also includes a scanning function block 65. In some examples, the scanning function block is implemented by a plurality of predetermined instructions stored on the memory 80, which, when executed by the processor 70, cause the processor 70 to perform the functions of the scanning function block 65.
[0052] In some examples, the scanning function block 65 and / or the fuzzer evaluation function block 50 scans the binary executable file 110 and generates: a list of points of interest; addresses of opcodes, each of which precedes a corresponding point of interest and is an opcode for a conditional check (i.e., a comparison of a variable with a predefined value); a list of strings of interest, such as service numbers, port numbers, keys, etc.; and a list of software stack characteristics, such as whether the stack is a transmission control protocol (TCP) stack, an Internet protocol (IP) stack, a crypto library, etc. Note that not all of the above information needs to be generated, and the scanning function block 65 and / or the fuzzer evaluation function block 50 may generate only some of the above information without exceeding the scope of the present disclosure.
[0053] In some examples, the scanning function block 65 and / or the fuzzer evaluation function block 50 generates a fuzzer agent 40, which includes the information generated above and also includes: code that allows hooks 111 to be added to the binary executable file 110 during runtime; code that sends information to a predetermined destination, which is external to the binary executable file 110 or within the binary executable file 110; one or more buffers for storing information of events; and code that optionally performs statistical checks and security checks, such as memory checks, function call monitoring, etc.
[0054] In some examples, upon initialization, one or more hooks 111 extract information from a memory stack associated with a corresponding point of interest. For example, information is extracted by using a pointer to an associated function and by extracting data from the memory stack starting at the address pointed to by the pointer, which points to data that needs to be read in order to enter the function. In such examples, the amount of memory read is determined based on a defined length that the function must read from the memory. In some examples, a bind function is used to read information from the memory stack. In some examples, the information includes an Internet Protocol (IP) address and a port number associated with the corresponding point of interest. The information is then used to generate a data unit so that the data unit arrives at the corresponding point of interest. As will be described below, reading information from the memory stack can also be performed after initialization.
[0055] The fuzzer data generator 20 generates data. In some examples, the fuzzer data generator 20 continuously generates data units. As used herein, the term "continuously" means that the fuzzer data generator 20 generates data units at predetermined time intervals within a predetermined time period. In some examples, the fuzzer data generator 20 generates at least 1,000 new data units (e.g., data packets) per second, and optionally generates at least 1 million new data units per second. As known to those skilled in the art of fuzz testing, the fuzzer data generator of the fuzzer (e.g., the fuzzer data generator 20) provides random input to the software in order to test the software or program. The input generated by the fuzzer data generator 20 can take various forms, such as a network packet, a file in a certain format, direct user input, a value, etc. In some examples, the fuzzer evaluation function block 50 controls the fuzzer data generator 20 to update the generated data unit at each time interval, so that the data unit generated at one time interval is different from the data unit generated at the next time interval.
[0056] In some examples, the fuzzer data generator 20 generates data according to predetermined rules. In some examples, the predetermined rules include information about the range of the following items: a memory address, a predetermined IP address, a predetermined port number, and / or a selected ECU, the range being defined as an area to be fuzz tested. In such examples, the target address of the generated data is set according to predetermined rules. In some examples, this information is extracted by the fuzzer data generator 20 and / or the fuzzer evaluation function block 50 according to a configuration file (such as a network communication description (NCD) file) and / or using an ECU extraction file. As described below, during runtime, the fuzzer evaluation function block 50 can identify changes in the address, port, and / or ECU that is the target. In such examples, the predetermined rules can be adjusted accordingly.
[0057] In some examples, the fuzzer evaluation function block 50 determines the predetermined rules based on a threat analysis and risk assessment (TARA). The fuzzer evaluation function block 50 may receive the TARA from an external device / network and / or from a user input terminal as known to those skilled in the art.
[0058] In some examples, the generated data unit is input into the DUT 115 by the input subsystem 30. As described above, for network-level fuzz testing, the input subsystem 30 inputs the generated data unit at the entry point of the process. For function-level fuzz testing, as described above, the input subsystem 30 directly inputs the generated data unit into the corresponding function. In some examples, the fuzz testing agent 40 communicates with the input subsystem 30, and each time the input subsystem 30 inputs a data unit into the DUT 115, the timestamp generator 60 of the fuzz testing agent 40 generates a corresponding timestamp. In such an example, when the hook 111 is reached, the timestamp generator 60 generates a corresponding timestamp. As used herein, the term "reached" means that the data flow has activated the corresponding hook 111.
[0059] In some examples, in response to the input data unit, each hook 111 outputs information associated with the corresponding point of interest to the fuzz testing agent 40. In particular, the corresponding point of interest is the point of interest to which the corresponding hook 111 is added. In some examples, as described below, the information includes data stored in an address of a memory (such as memory 80) associated with the corresponding predetermined point (e.g., a value stored in a memory address range pointed to by a pointer to a corresponding function, a value of a pointer to a corresponding function, a corresponding IP number, and / or a corresponding port number). In some examples, the information associated with the corresponding point of interest indicates a security vulnerability of the DUT 115. In some examples, the information associated with the corresponding point of interest includes an indication of a security vulnerability associated with a heap or stack associated with the executable binary file 110. In another example, alternatively or additionally, the information associated with the corresponding point of interest includes an indication of a library access. In some examples, the information associated with the corresponding point of interest includes an indication of a memory stack overflow or a memory heap overflow. This may include pointing to an address outside the address range of the memory stack or memory heap.
[0060] In some examples, alternatively or additionally, the information associated with the corresponding point of interest includes an indication of arrival at the corresponding point of interest. In such an example, the fuzzer evaluation function block 50 performs a statistical evaluation of the number of times each predetermined point of interest is initiated. The results of the statistical analysis are compared with predetermined parameters and thresholds to determine whether there is a security vulnerability.
[0061] In some examples, as described above, the information associated with the corresponding point of interest may also include data copied from the memory stack. In some examples, as described above, the IP address and / or port number associated with the corresponding point of interest is read. In some examples, the fuzzer evaluation function block 50 compares the information copied from the memory stack with the corresponding information copied from the memory stack at initialization. If there is a difference in this information, such as a change in the IP address or port number, the fuzzer evaluation function block 50 outputs an indication of the existence of such a difference. In some examples, such an indication is added to a report indicating a security vulnerability and / or software error (bug) present in the DUT 115.
[0062] In some examples, the fuzzer evaluation function block 50 evaluates the received information to identify problems in control flow integrity (CFI). For example, the fuzzer evaluation function block 50 compares the value of the pointer of the corresponding function with the stored address value associated with the corresponding function. If the value of the pointer is not equal to the stored address value, the fuzzer evaluation function block 50 determines that there is a problem with the CFI, and in some examples outputs an indication that there is such a problem, which optionally includes the value of the pointer and information about the corresponding data unit of the input.
[0063] In some examples, information associated with the corresponding point of interest is stored in a predetermined portion of the global buffer. In some examples, each portion of the global buffer is associated with a corresponding hook 111. In some examples, an identifier for each task that may include a corresponding hook 111 is stored in each portion of the global buffer. In some examples, a dedicated debug unified diagnostics service (UDS) data identifier (DID) is used to read information in the global buffer. In another example, an existing UDSDID is used to read the global buffer. In some examples, data is read from the buffer via the UDSDID using a diagnostic communication manager (DCM) call or a DCM service port.
[0064] In another example, the fuzz testing agent 40 is configured to send information to the fuzzer evaluation functional block 50 using a User Datagram Protocol (UDP), Controller Area Network (CAN) message. In some examples, the fuzz testing agent 40 sends one or more data packets with the information to the fuzzer evaluation functional block 50. In some examples, the fuzz testing agent 40 sends multiple copies of the information to the fuzzer evaluation functional block 50. In another example, the fuzz testing agent 40 additionally sends one or more cookies along with the data so that the fuzzer evaluation functional block 50 can track whether any data from the fuzz testing agent 40 has not arrived.
[0065] In another example, the debugger continuously polls the global buffer, and optionally, the read data is output to the testing tool 100 via an application interface (eg, a Windows dll file).
[0066] In some examples, in response to the corresponding output information of the hook 111, the fuzz testing agent 40 determines which of the input data units arrives at the corresponding hook 111. In some examples, the timestamp generator 60 generates a timestamp when each data unit is input by the input subsystem 30, and when each hook 111 is reached, in which case, determining which of the input data units arrives at the corresponding hook 111 is responsive to the generated timestamp. In particular, the fuzz testing agent 40 compares the timestamp generated when the corresponding hook 111 is reached with the timestamp generated when the data unit is input. The difference between the timestamps is compared with a predetermined time lapse threshold, and in response to one of the differences being within a predetermined range of the time lapse threshold, the associated data unit is determined to be a data unit that arrives at the corresponding hook 111. In another example, determining which of the input data units arrives at the corresponding hook 111 is performed by the fuzzer evaluation function block 50.
[0067] In some examples, a dedicated counter is provided for each point of interest. The counter may be implemented in any of: a corresponding hook 111; a fuzz testing agent 40; and a fuzzer evaluation function 50. The counter indicates the number of times the point of interest has been reached. This information may be used for statistical analysis (as described above), and for updating data units (as will be described below).
[0068] The fuzzer data generator 20 is responsive to the output of the fuzzer evaluation function block 50. In some examples, the fuzzer evaluation function block 50 indicates to the fuzzer data generator 20 how the data unit should be updated (e.g., which bits of the data unit should be transformed to perform the fuzz testing process). In another example, the fuzzer evaluation function block 50 controls the fuzzer data generator 20 to update the data unit. In some examples, a selected portion of the data unit is updated randomly. In another example, a selected portion of the data unit is updated according to a predetermined rule or model. In another example, a selected portion of the data unit is updated in response to a detected security vulnerability.
[0069] In some examples, the fuzzer data generator 20 generates a data unit in response to a result of determining which of the input data units arrives at the corresponding hook 111. In particular, if a specific data unit arrives at the corresponding hook 111, the fuzzer evaluation function block 50 causes the fuzzer data generator 20 to generate an updated data unit using the specific data unit as a reference. Advantageously, the information received by the hook 111 allows the data unit input to the DUT 115 to be updated more efficiently.
[0070] In some examples, where the fuzzer evaluation function block 50 and the fuzzer data generator 20 are embedded in the binary executable file 110 , the fuzzer evaluation function block 50 controls the fuzzer data generator 20 to input the data unit directly into the corresponding function of the binary executable file 110 .
[0071] In some examples, the input data unit is continuously updated until each hook 111 is reached. In another example, the input data unit is continuously updated until each hook 111 is reached at least a predetermined number of times. In some examples, the evaluation function block 50 generates multiple instances of attack scenarios, and for each batch of scenarios, there is a corresponding subset of hooks 111 added to the binary executable file 110. Advantageously, the performance impact of the hooks 111 is negligible, and maximum coverage is achieved after repeatedly running all scenarios.
[0072] In some examples, in response to the information received at the fuzzer evaluation function block 50, the fuzzer evaluation function block 50 outputs an indication of the corresponding point of interest to the fuzz testing agent 40. In response to the output indication of the corresponding point of interest, the fuzz testing agent 40 adds the corresponding hook 111 to an additional location in the binary executable file 110 associated with the corresponding point of interest. In some examples, the additional location is located earlier in the flow of the binary executable file than the corresponding point of interest. As used herein, the term "earlier in the flow" means that the instructions of the additional location are executed before the instructions of the corresponding point of interest.
[0073] In some examples, the fuzzer evaluation function block 50 outputs an indication of the corresponding point of interest to the fuzz testing agent 40 in response to not receiving information associated with reaching the corresponding point of interest within a predetermined amount of time interval. In particular, if the corresponding hook 111 has not been reached after a predetermined number of data units have been input, the fuzz testing agent 40 adds another hook 111 at an earlier point in the process. In some examples, the additional hook 111 can be added in response to analyzing the stack to determine which points in the binary executable file 110 are being affected by the input data units.
[0074] In some examples, in response to a corresponding hook 111 of the one or more hooks 111 not being activated within a corresponding predetermined amount in a predetermined time interval, the fuzzer evaluation function block 50 identifies a comparison opcode located before the corresponding hook 111. In some examples, the comparison opcode is found by searching the assembly code of the first comparison instruction before the corresponding hook 111. The comparison opcode is associated with one or more comparison values and one or more variable values (stored in a dedicated register). In particular, the comparison can be performed between several registers and corresponding values. The following is described with respect to a single variable value and a single comparison value, however, this is not meant to be limiting in any way.
[0075] As used herein, the term "variable value" refers to a value of a variable that is not constant. As used herein, the term "comparison value" refers to a predetermined value used to compare with the variable value. If the variable value is equal to the comparison value, the comparison condition is satisfied.
[0076] The fuzzer evaluation function block 50 repeatedly receives the comparison value and the variable value of the comparison instruction from the fuzz testing agent 40 at multiple instances of a predetermined time interval. In some examples, the corresponding hook 111 includes a wrapper function that reads the variable value and the comparison value from the memory, and the branch instruction of the corresponding hook 111 includes the read value.
[0077] In some examples, when the fuzzer evaluation function block 50 reads the variable value from the register, at least a predetermined number of data units are input. Additionally, the fuzzer evaluation function block 50 controls the fuzzer data generator 20 to repeatedly adjust the generated data units in response to the comparison value and the variable value. In particular, the generated data units are adjusted so that the variable value will be equal to the comparison value. In some examples, the variable value is stored by the fuzzer evaluation function block 50 for each time interval.
[0078] In response to the variable value being equal to the comparison value, the fuzzifier evaluation function block 50 determines the necessary adjustments to the generated data unit to make the variable value equal to the comparison value. For example, the fuzzifier evaluation function block 50 determines which bits of the data unit need to be adjusted to which values in order to meet the comparison condition and thus reach the corresponding hook 111, as will be described below. The fuzzifier evaluation function block 50 then controls or instructs the fuzzifier data generator 20 what adjustments need to be made to the data unit to meet the comparison condition.
[0079] In some examples, the generated data unit is repeatedly adjusted in response to a predetermined optimization algorithm until the variable value is equal to the comparison value. In particular, the optimization algorithm adjusts the input data unit and follows the variable value until it becomes equal to the comparison value. In some examples, the predetermined optimization algorithm is a gradient descent algorithm. In particular, as known to those skilled in the art, the gradient descent algorithm is a first-order iterative optimization algorithm for finding a local minimum of a differentiable function.
[0080] Sometimes, a function is only reached under rare circumstances. As an example, the function memcpy can be inside an if condition, such as the following:
[0081] if(X
[100] == 'R'){
[0082] memcpy(a,b,c);
[0083] }
[0084] In such an example, memcpy will be reached very rarely. Advantageously, the above method allows fuzz testing of the function memcpy in a minimal period of time.
[0085] In some examples, before applying a predetermined optimization algorithm to determine the necessary adjustments, the fuzzer evaluation function block 50 is configured to repeatedly control or instruct the fuzzer data generator 20 to insert predetermined values into corresponding positions of corresponding data units, and the corresponding positions are different at each repetition. For example, in the first iteration, '$' can be inserted into all bytes of the data unit. Then, the fuzzer evaluation function block 50 analyzes the memory stack associated with the binary executable file 110 to determine which corresponding positions in the input data unit affect the memory stack. In the above example, the fuzzer evaluation function block 50 will analyze the stack to determine which address now contains '$'.
[0086] In some examples, the generated data units are repeatedly adjusted until the variable value is equal to the comparison value. In some examples, the adjustment is responsive to the result of determining the corresponding position. In particular, as described above, a particular section of each data unit is identified as affecting an address near the corresponding hook 111. In some examples, as described above, the section in each new data unit is changed until the variable value is equal to the comparison value. For example, if the identified section is the 10th byte of the payload of the data unit, the 10th byte of each new data unit is adjusted until the variable value is equal to the comparison value.
[0087] In some examples, such as Figure 2AAs shown, the fuzz testing system 10 also includes a machine learning (ML) subsystem 200. As described above, in some examples, the ML subsystem 200 is implemented by instructions stored on a memory and executed by one or more processors. In another example, all or part of the ML subsystem 200 is implemented on a network (such as a cloud-based network). In some examples, the ML subsystem 200 includes: one or more convolutional neural network (CNN) trainers 203; and one or more CNNs 205. As used herein, the term "CNN trainer" refers to a system or software instruction set running on a processor to train the corresponding CNN 205, as known to those skilled in the art. In particular, the CNN trainer trains the CNN by passing the input through the CNN and comparing the output with acceptable parameters / values. In some examples, the training includes: a forward phase, in which the input is passed through the network in its entirety; and a backward phase, in which the gradient is back-propagated and the weights are updated. As known to those skilled in the art, "back propagation" is short for error back propagation, which is an algorithm for supervised learning of artificial neural networks using gradient descent.
[0088] like Figure 2B As shown, in some examples, the subsystem 200 also includes a data unit functional block 210. Although four CNNs 205 and four CNN trainers 203 are shown, this is not meant to be limiting in any way, and any number of CNNs 205 and CNN trainers 203 (including one) may be provided without exceeding the scope of the present disclosure. In some examples, the data unit functional block 210 communicates with the fuzzifier evaluation functional block 50 via a network interface or other suitable communication means.
[0089] In some examples, the fuzzer evaluation function block 50 is configured to store corresponding variable values at predetermined time intervals. The CNN trainer 203 of the ML subsystem 200 trains the CNN 205 using the above-mentioned stored variable values and the corresponding generated data units associated with the stored variable values. In particular, for each data unit, there is a corresponding variable value, which appears in a register, and one or more CNNs 205 are trained using the variable value and the corresponding data unit. In some examples, as shown in the figure, multiple CNNs 205 are trained in parallel. In some examples, the training is performed using a binary cross entropy loss function.
[0090] In some examples, if the loss function of at least one CNN 205 converges to a sufficiently low predetermined value, the corresponding CNN 205 will include a model that receives a data unit and outputs a value indicating what the variable value would be if the corresponding data unit was input into the DUT 115.
[0091] Figure 3A A high-level block diagram of a system 215 for fuzz testing according to some examples is shown. System 215 is similar to system 10 in all respects, but adds ML subsystem 200, simulator 115', and input subsystem 30'. In some examples, input subsystem 30' includes instructions that, when read by one or more processors, cause input subsystem 30' to access various functions of a process running in simulator 115'. In some examples, simulator 115' includes a virtual machine or other virtual environment (optionally running in a cloud computing environment) that simulates DUT 115. In some examples, as known to those skilled in the art, simulator 115' includes input and output for simulating ports and CPUs of DUT 115. Simulator 115' includes a copy 110' of a binary executable file 110 and a fuzz testing agent 40' embedded in the copy binary file 110'. Input subsystem 30' directs data to one or more functions within copy 110' of binary executable file 110. In some examples, as will be described below, fuzz testing agent 40' may be different than fuzz testing agent 40. Fuzz testing agent 40' is implemented by a plurality of instructions stored on a memory that, when executed by one or more processors, cause the one or more processors to perform the functionality of fuzz testing agent 40'.
[0092] In some examples, the emulator 115 ′ is implemented by a plurality of instructions stored on a memory that, when read by one or more processors, cause the one or more processors to implement the functionality of the emulator 115 ′.
[0093] Figure 3A Only a single CNN trainer 203 and a single CNN 205 are shown, however, this is not meant to be limiting in any way, and any number of CNNs 205 and corresponding CNN trainers 203 may be provided without exceeding the scope. In some examples, the output of the fuzzifier evaluation function block 50 communicates with the input of each CNN trainer 203. Although Figure 3A A direct connection between the fuzzifier evaluation block 50 and the CNN trainer 203 is shown, but this is not meant to be limiting in any way. In another example (not shown), an additional system is provided to receive information from the fuzzifier evaluation block 50 and input information into the CNN trainer 203. As described above, each CNN trainer 203 trains a corresponding CNN 205, and in some examples, the output of the CNN 205 communicates with the input of the data unit block 210, and the output of the data unit block 210 communicates with the input of the fuzzifier evaluation block 50.
[0094] In some examples, the fuzzer data generator 20 is responsive to the output of one or more CNNs 205. In particular, in such an example, the data unit function block 210 sends a data unit verified by the CNN 205 as satisfying a condition (i.e., the output of the corresponding CNN 205 is equal to the comparison value) to the fuzzer evaluation function block 50. Then, the fuzzer evaluation function block 50 instructs the fuzzer data generator 20 to generate such a data unit for the input subsystem 30. Then, the fuzzer evaluation function block 50 analyzes whether the data unit can actually satisfy the condition and reach the point of interest. Although the above is described in the case where the data unit function block 210 sends the data unit to the fuzzer evaluation function block 50, this is not meant to be limiting in any way. In other examples, the data unit function block 210 may send the data unit to the fuzzer data generator 20 or the input subsystem 30 without exceeding the scope of the present disclosure. Therefore, in order to fuzz test the corresponding point of interest, there is a difficult condition before the point of interest, and the CNN 205 and the data unit function block 210 provide a data unit that satisfies the condition, thereby reaching the corresponding hook 111.
[0095] Then, the fuzzifier evaluation function block 50 receives the variable value associated with the input data unit, and in some examples, outputs an indication of whether the corresponding variable value is equal to the corresponding comparison value to the CNN trainer 203. In some examples, the indication includes a binary value, a Boolean value, or the like. In another example, the indication includes the corresponding variable value, and the fuzzifier evaluation function block 50 and / or the CNN trainer 203 determines whether the corresponding variable value is equal to the corresponding comparison value.
[0096] In some examples, when the corresponding variable value is equal to the corresponding comparison value, the training of CNN 205 is completed. When the corresponding variable value is not equal to the corresponding comparison value, which means that the model of CNN 205 is inaccurate, the CNN trainer 203 inputs the corresponding data unit sent to the input subsystem 30 into one or more CNNs 205 to continue its training, that is, the training of CNN 205.
[0097] In some examples, in response to the fuzzifier evaluation function block 50 indicating that the data unit successfully reaches the point of interest, the CNN trainer 203' trains the second CNN 205' to generate data units that have a high chance of reaching the point of interest based on the successful data units, such as Figure 3BAs shown. In particular, the successful data units provided by the fuzzifier evaluation functional block 50 and / or the data unit functional block 210 are used by the CNN trainer 203' (optionally, the CNN trainer 203' is one or more CNN trainers 203) to train the CNN 205' so that the trained CNN 205' generates data units that satisfy the conditions at the points of interest. Therefore, in such an example, the data units generated by the trained CNN 205' are sent to the input subsystem 30 or the fuzzifier data generator 20 for input into the DUT 115.
[0098] In some examples, in the case where the CNN 205 does not converge correctly, the fuzzer evaluation function block 50 obtains a snapshot of the memory stack / heap associated with the binary executable file 110 and the registers of the CPU memory. As used herein, the term "snapshot" refers to the instructions and values stored in each address (including the CPU memory registers) from the beginning of the process to the corresponding point of interest (e.g., memcpy). In response to the snapshot, the fuzzer evaluation function block 50 uses the snapshot to set the memory of the simulator 115' to have the same value and state as the memory of the CPU at the snapshot time (when the binary file (binary) 110 is running in the DUT 115). In the simulator 115', the fuzzer evaluation function block 50 optionally uses the CNN to insert various values into various variables of the function containing the corresponding point of interest (e.g., the function containing memcpy and the corresponding condition) until the variable value is equal to the comparison value.
[0099] The above may be used, among other things, to: generate rule sets for firewalls; coverage reports (ie, how much of the DUT 115 was tested); and security vulnerability statistics.
[0100] Figure 3C A diagram depicting an example of a first operational flow of the system 215 for fuzz testing is shown. In step A1, the fuzz testing agent 40 sends initialization information to the fuzzer evaluation functional block 50. As described above, in some examples, the initialization information included by the fuzz testing agent 40 is provided by the scanning functional block 65. In step A2, the fuzzer evaluation functional block 50 instructs the fuzz testing agent 40 to add a hook 111 to the process of the binary executable file 110 during runtime.
[0101] In step A3, the fuzzer evaluation function block 50 updates the fuzzer data generator 20 regarding which bits of each data unit are to be modified during the fuzz testing process. In particular, as known to those skilled in the art, during fuzz testing, data units are constantly modified in order to test the system or parts thereof. Therefore, the fuzzer evaluation function block 50 determines which parts of the data unit need to be modified to perform the fuzz testing process. The parts may be determined based on: the location of the point of interest being fuzz tested, e.g., a portion of the data unit that affects the point of interest; an address defined in the initialization information as being within the address space of the process; and / or other relevant parameters.
[0102] In step A4 , the fuzzifier data generator 20 generates a data unit based on the information received from the fuzzifier evaluation function block 50 , and sends the generated data unit to the input subsystem 30 , and the data unit is then input into the DUT 115 .
[0103] In step A5, when the process flow reaches the hook, the fuzz testing agent 40 sends event information associated with the corresponding hook 111 to the fuzzer evaluation function block 50. In response to the received information, the fuzzer evaluation function block 50 updates the fuzzer data generator 20. As will be described below, in some examples, the event information may include: information about POI events, i.e., notifications that the corresponding point of interest has been reached; information about coverage events, i.e., notifications that the corresponding code block has been reached; CFI events, i.e., notifications of problems in the control flow, such as detection of a crash, memory corruption, incorrect flow, etc.; and / or information about statistical events, i.e., counts of the number of times the corresponding hook has been reached or process-level statistics, such as average CPU load, free stack available memory, the number of page fault interrupts in one second, etc.
[0104] Thus, for example, in response to information about a POI event or a coverage event, the fuzzer evaluation function block 50 may instruct the fuzzer data generator 20 to maintain the values in a portion of the data unit that caused the process to reach the point of interest / code block, and to modify other portions of the data unit for fuzz testing purposes. In some examples, in response to information about a CFI event, the fuzzer evaluation function block 50 may instruct the fuzzer data generator 20 to update a predetermined portion of the data unit so as to face a different point of interest.
[0105] In some examples, in response to information about statistical events, the fuzzer evaluation function block 50 may instruct the fuzzer data generator 20 to change corresponding portions of the data unit in order to continue the fuzz testing process. For example, if an abnormal statistical event is detected, the fuzzer evaluation function block 50 updates the instruction set / model for modifying the data unit so that further statistical events will be caused, and the fuzzer evaluation function block 50 instructs the fuzzer data generator 20 to modify the data unit accordingly.
[0106] Figure 3D A diagram is shown describing an example of a second operational flow of the system 215 for fuzz testing using a CNN model to overcome conditional checks. In some examples, Figure 3D The second process is Figure 3C The second process may be an extension of the first process, however, the second process may also be separate from the first process.
[0107] In step B1, the fuzzer evaluation function block 50 sends an instruction to the fuzz testing agent 40 to add a hook 111 on the condition check that is closest to the corresponding point of interest (i.e., the condition that is checked in order to allow the process to reach the point of interest). The closest condition check is defined as the first condition check before the corresponding point of interest. Note that the closest condition check does not have to be immediately before the point of interest, and there may be one or more instructions between the condition check and the corresponding point of interest. As described above, in some examples, adding the hook 111 includes replacing the opcode of the condition check with a branch instruction to the fuzz testing agent 40.
[0108] In some examples, fuzzer evaluation function 50 communicates with fuzz testing agent 40 by instructing fuzzer data generator 20 to generate a data unit for fuzz testing agent 40. For example, the data unit may be a UDP packet whose header contains the IP address and / or port of fuzz testing agent 40.
[0109] In step B2, when the process reaches the hook 111 of step B1, the fuzz testing agent 40 sends event information associated with the corresponding hook 111 to the fuzzer evaluation function block 50, as described above in conjunction with step A5.
[0110] In step B3, in response to the information received in step B2, the fuzzer evaluation function block 50 sends relevant information to the CNN trainer 203, which includes: a data unit that enables the process to reach the hook 111, which is optionally identified by a timestamp generated at the input subsystem 30 and the corresponding hook 111; and corresponding register values, including comparison values and variable values as described above.
[0111] In step B4, the CNN trainer 203 uses the bits in the data unit bits as the input layer and the register values as the output layer to train the CNN model. When the model converges, the model is sent to the data unit function block 210. As described above, in some examples, multiple CNN trainers 203 run in parallel.
[0112] In step B5, the data unit functional block 210 runs the model in several parallel instances within the computing environment (e.g., in a cloud computing environment), using random input bits for each instance. In response to reaching the desired output, i.e., a data unit that causes the output variable value of the model to be equal to the comparison value, the input bits are sent to the fuzzifier evaluation functional block as data unit candidates.
[0113] In step B6, the fuzzifier evaluation function block 50 instructs the fuzzifier data generator 20 to send the data unit candidate to the input subsystem 30. In step B7, the fuzzifier data generator 20 sends the data unit candidate to the input subsystem 30.
[0114] In step B8, when the process flow reaches the hook 111 of steps B1 and B2, the fuzz test agent 40 sends event information (including variable values) associated with the corresponding hook 111 to the fuzzer evaluation function block 50, as described above. In some examples, the fuzzer evaluation function block 50 compares the variable value with the comparison value, and if the condition is met, the fuzz test agent 40 branches to the next opcode to continue the process flow until the corresponding point of interest is reached. In this case, the data unit candidate is defined by the fuzzer evaluation function block 50 as a verified data unit, and the verified data unit is used as the basis for subsequent iterations of data units for reaching the next block or point of interest.
[0115] In step B9, the fuzzifier evaluation function block 50 sends the verified data unit to the CNN trainer 203' to train the CNN model 205' to generate a data unit similar to the verified data unit, that is, a data unit that generates the same conditions to overcome the conditional check. In the case where the data unit candidate does not overcome the conditional check, the fuzzifier evaluation function block 50 sends the variable value realized by the data unit candidate to the CNN trainer 203, and the CNN trainer 203 uses this information to continue training the CNN model 205.
[0116] Figure 3E A diagram is shown describing an example of a third operational flow of a system 215 for fuzz testing that uses a CNN model to perform function-level fuzz testing. In some examples, Figure 3E The third process is Figure 3C The first process and / or Figure 3DThe third process is an extension of the second process, however, the third process can also be separated from the first process and the second process.
[0117] In step C1, the fuzzer evaluation function block 50 sends an instruction to the fuzz test agent 40 to add a hook 111 at the entry point of a predetermined function. In step C2, the fuzzer data generator 20 sends a data unit to the input subsystem 30, and then the input subsystem 30 inputs the data unit into the DUT 115. As described above, the data unit is generated to target the corresponding function.
[0118] In step C3 , as described above in conjunction with steps A5 and B2 , when the process flow reaches the corresponding hook 111 , the fuzz testing agent 40 sends event information associated with the corresponding hook 111 to the fuzzer evaluation function block 50 .
[0119] In step C4, upon receiving the event information, the fuzzer evaluation function block 50 sends the memory snapshot to the fuzz test agent 40'. In some examples, as described above in conjunction with the communication between the fuzzer evaluation function block 50 and the fuzz test agent 40, the memory snapshot is sent to the fuzz test agent 40' via the input subsystem 30'. In another example, the fuzzer evaluation function block 50 communicates directly with the fuzz test agent 40'. As will be further described below, in response to the received snapshot, the fuzz test agent 40' initiates function-level fuzz testing within the emulator 115'. In some examples, in response to the received memory snapshot, the fuzz test agent 40' sets the corresponding value of the emulator 115' to the corresponding value of the DUT 115, so that the data unit input at the input subsystem 30' will arrive at the corresponding function of step C1. In some examples, the set value includes a register value from the memory snapshot. In some examples where the emulator 115' is a QEMU emulator, a protocol such as the QEMU machine protocol (QMP) is used to set the register value.
[0120] In step C5, the fuzzer evaluation function block 50 controls the fuzzer data generator 20 to generate a data unit and sends it to the input subsystem 30', and in step C6, the fuzzer data generator 20 generates a data unit and sends it to the input subsystem 30'. The generated data unit is intended to perform fuzz testing on the corresponding function, that is, the relevant part of the data unit is continuously modified to perform fuzz testing on the corresponding function.
[0121] In step C7 , as will be further described below, the fuzz testing agent 40 ′ sends event information associated with the corresponding function to the fuzzer evaluation functional block.
[0122] In some examples, the evaluation function 50 generates one or more reports on CFI events and statistical events. The generated reports can be stored in a database and / or transmitted to an external system / server.
[0123] Figure 4A A high-level block diagram of an example of a system 300 for fuzz testing is shown, and Figure 4B A high-level block diagram of a more detailed example of a system 300 for fuzz testing is shown.
[0124] In some examples, system 300 includes: fuzzer data generator 20; input subsystem 30; fuzz testing agent 40 embedded in binary executable file 110, which is initialized to run on DUT 115; fuzzer evaluation function block 50; reporting function block 130; and memory 140. Although not shown, various timestamp generators can be provided, as described above in conjunction with system 10. Fuzz testing agent 40 is implemented as described above, however, Figure 4A An example is shown in which the fuzz testing agent 40 includes an event handler 41 and a network manager 42. In some examples, such as Figure 4G As shown, the fuzz testing agent 40 includes a plurality of event handlers 41. Although three event handlers 41 are shown, this is not meant to be limiting in any way, and in another example, any number of event handlers 41 may be provided without exceeding the scope of the present disclosure.
[0125] The fuzzer evaluation function block 50 is implemented as described above, however, Figure 4A An example is shown in which the fuzzer evaluation function block 50 includes a fuzz testing unit 51 and a control unit 52 .
[0126] In some examples, the event handler 41 is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, causes the one or more processors to perform the functions of the event handler 41. In some examples, the one or more processors are implemented as part of the DUT 115. In some examples, the network manager 42 is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, causes the one or more processors to perform the functions of the network manager 42. In some examples, the one or more processors are implemented as part of the DUT 115. In some examples, the event handler 41 and the network manager 42 are implemented on the same one or more processors. In some examples, the network manager 42 implements a UDP server configured to listen on one or more predetermined ports.
[0127] In some examples, the fuzz testing unit 51 is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, causes the one or more processors to perform the functions of the fuzz testing unit 51. In some examples, the control unit 52 is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, causes the one or more processors to perform the functions of the control unit 52.
[0128] In some examples, the reporting function block 130 is implemented by a plurality of instructions stored on a memory (optionally a memory 140), which when executed by one or more processors cause the one or more processors to perform the functions of the reporting function block 130. In some examples, the reporting function block 130 communicates with an external system or server. In some examples, the reporting function block 130 includes a memory or communicates with a memory 140.
[0129] In some examples, memory 140 (and similarly, memory 80 described above) includes persistent memory, i.e., non-volatile memory, such as a solid-state drive (SSD), a NAND flash drive, a ferroelectric RAM, etc. In some examples, memory 140 (and similarly, memory 80 described above) is implemented as a corresponding portion of a memory for DUT 115.
[0130] In some examples, such as Figure 4B As shown, the system 300 for fuzz testing also includes: an emulator 115'; a fuzz testing agent 40'; a fuzz testing unit 51'; and a control unit 52'. The fuzz testing agent 40' includes an event handler 41' and a network manager 42'. A copy 110' of the binary executable file 110 is implemented on the emulator 115'.
[0131] In some examples, the event handler 41' is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, cause the one or more processors to perform the functions of the event handler 41'. In some examples, the network manager 42' is implemented by a plurality of instructions stored on a memory (optionally, the memory 140), which, when executed by one or more processors, cause the one or more processors to perform the functions of the network manager 42'. In some examples, as described below, the network manager 42' may include a network socket configured for network communications.
[0132] In some examples, the fuzz testing unit 51' is implemented by a plurality of instructions stored on a memory (optionally the memory 140), which, when executed by one or more processors, cause the one or more processors to perform the functions of the fuzz testing unit 51'. In some examples, the control unit 52' is implemented by a plurality of instructions stored on a memory (optionally the memory 140), which, when executed by one or more processors, cause the one or more processors to perform the functions of the control unit 52'. In some examples, the event handler 41', the network manager 42', the fuzz testing unit 51', and the control unit 52' are all implemented by the same one or more processors that implement the emulator 115'.
[0133] In some examples, event handler 41' is embedded within binary copy 110', while network manager 42', fuzz testing unit 51', and control unit 52' are implemented within emulator 115' but are not embedded within binary copy 110'. In some examples, network manager 42' communicates with event handler 41' using memory shared between the two processes.
[0134] Although the system 215 and the system 300 are described in the examples as including one or more emulators 115', this is not meant to be limiting in any way. Alternatively or additionally, the system 215 and / or the system 300 may include one or more virtual machines, such as AWS Graviton servers commercially available from Amazon Web Services. In the event that the binary file 110 calls a function that is not supported by the virtual machine, the function may be replaced with a compatible function that simulates the operation of the original function.
[0135] Figure 4C A high-level flow chart of an example of a method for fuzz testing is shown. In some examples, the described fuzz testing method is implemented using system 300, however, this is not meant to be limiting in any way. In step 400, the binary executable file 110 is analyzed to determine relevant information. As described above, the analysis can include identifying: a list of points of interest; the addresses of opcodes, each of which precedes a corresponding point of interest and is an opcode for a conditional check (i.e., a comparison of a variable with a predefined value); a list of strings of interest, such as service numbers, port numbers, keys, etc.; and a list of software stack characteristics, such as whether the stack is a transmission control protocol (TCP) stack, an Internet protocol (IP) stack, a crypto library, etc. In some examples, as described above, the analysis is performed by a scanning functional block 65 (not shown for simplicity). In another example (not shown), as described above, the analysis is performed by a fuzzer evaluation functional block 50, in particular, by a control unit 52.
[0136] In some examples, as described above, the binary executable file 110 is analyzed to define points of interest. As described above, in some examples, the defined points of interest are functions of a predetermined type. In another example, alternatively or additionally, an indication of a point of interest is received from a user input.
[0137] In some examples, the binary executable file 110 is analyzed to identify a block graph for each point of interest. In particular, if there are one or more code blocks leading to the corresponding point of interest, these code blocks are identified. For example, Figure 4F As shown, FUNC2 is a function defined as a point of interest. As shown, to reach FUNC2, the process starts at block_0x092 and passes through block_0x099 and block_0x122 until block_0x111 containing FUNC2 is reached. As used herein, the term "code block" refers to multiple lines of code grouped together. The numbers shown (0x092, 0x099, 0x122, and 0x111) indicate the memory address of the first opcode in the code block. In some examples, a code block is defined as a plurality of instructions that start with a branch instruction and end with a branch instruction.
[0138] In some examples, certain metadata (eg, certain strings) are identified within binary executable file 110 .
[0139] The binary executable 110 is instrumented to add it to the DUT 115. As described above, in some examples, the fuzz testing agent 40 is embedded within the instrumented binary. As described above, in some examples, the fuzz testing agent 40 includes: code for implementing the network manager 42, optionally code for sending and receiving UDP packets, i.e., code; hooks inserted into the binary executable 110 at initialization; code for implementing the event manager 41 and optionally storing information; code for adding hooks during runtime; or any combination of the above options.
[0140] In some examples, user input is received at the fuzzer evaluation functional block 50, the user input defining: the number of data units to be sent for each point of interest; and / or the maximum time allowed for fuzz testing each point of interest. In some examples, user input is received at the fuzzer evaluation functional block 50, the user input defining traffic configuration information related to the traffic policy allowed to the DUT 115. In some examples, user input is received at the fuzzer evaluation functional block 50, the user input including TARA information about the binary executable file 110. Any one or a combination of the above user inputs may be received at the fuzzer evaluation functional block 50.
[0141] In step 410, in phase 1, network-level fuzz testing is performed, as will be described below. In step 420, in phase 2, as will be described below, when the process flow reaches a point of interest, the point of interest is fuzz tested using function-level fuzz testing. In response to detecting a CFI event in the function-level fuzz testing of phase 2 (step 420), the probability that the CFI event actually occurred is checked in the following two steps: step 430, using network-level fuzz testing, as will be described below; and step 440, using function-level fuzz testing, as will be described below.
[0142] Figure 4D A high level flow chart showing the flow of a portion of the operation of the fuzz testing method. The method is described with respect to system 300, however, this is not meant to be limiting in any way. In step 500, the binary executable file 110 is analyzed as described above in conjunction with step 400.
[0143] In step 510, the control unit 52 or scanning function block 65 (not shown) of the fuzzer evaluation function block 50 determines whether the binary executable file 110 is new or whether it has been previously fuzz tested by the system 300. In the event that it is determined that the binary executable file 110 is not new (i.e., it has been previously fuzz tested by the system 300), data is extracted from the memory 140 and / or the reporting function block 130 regarding: previous coverage reports, such as reports on which points of interest were previously reached and how often they were reached; and / or scenarios for reaching a particular point of interest, such as reports on data units that successfully reached the corresponding point of interest.
[0144] In step 520, in some examples, a list of new points of interest is generated based on a comparison of the analysis of step 500 with the results of one or more previous fuzz testing sessions. In particular, in some examples, points of interest that have not been reached in previous fuzz testing sessions are defined. In another example, both new points of interest and previously fuzz tested points of interest are defined in the list.
[0145] In step 540, using the information of step 510, the control unit 52 of the fuzzer evaluation functional block 50 instructs the fuzz test unit 51 to execute: control the fuzzer data generator 20 to use the data unit that successfully reached certain points of interest before. Then, the control unit 52 determines the coverage report of the fuzz test, that is, how many defined points of interest are reached and how many times they are reached. In step 550, the currently determined coverage report is compared with one or more previous coverage reports.
[0146] In step 560, it is determined whether the coverage reports are the same. If the result of the comparison indicates that the coverage reports are the same, or the difference is less than one or more predetermined thresholds, then in step 570, the control unit 52 instructs the fuzzy testing unit 51 to perform fuzzy testing on the new interest points (i.e., interest points that have not been fuzzy tested before).
[0147] If the comparison result of step 550 indicates that the coverage reports are not the same, or the difference is not less than one or more predetermined thresholds, then in step 580, the control unit 52 instructs the fuzz testing unit 51 to fuzz test all the points of interest in the list (including the points of interest that have been fuzz tested previously) again. Similarly, if the comparison result of step 510 indicates that the binary executable file 110 has not been fuzz tested before, the points of interest are not skipped.
[0148] In step 590, after the fuzz testing of step 570 and / or step 580, the control unit 52 controls the reporting function block 130 to store information about the fuzz testing session, which information optionally includes: an identifier of the binary executable file 110; a coverage report determined by the control unit 52; a scenario of reaching the corresponding point of interest, i.e., a specific data unit reaching the corresponding point of interest; or any combination thereof.
[0149] Figure 4E A high-level flow chart of the flow of a portion of the operation of the fuzz testing method is shown. The method is described with respect to system 300, however, this is not meant to be limiting in any way. In step 600, for each defined point of interest, a corresponding hook is placed at the point of interest, as described above in conjunction with step 410. In some examples, a hook is also added at the beginning of each code block in the call tree of the corresponding point of interest. In some examples, each hook is added by the fuzz testing agent 40. In another example, one or more hooks are added by the control unit 52 of the fuzzer evaluation functional block 50 and / or the scanning functional block 65.
[0150] In some examples, where the fuzzer evaluation function block 50 instructs the fuzz testing agent 40 to add a hook, the fuzzer evaluation function block 50 sends a message (such as a UDP message) to the fuzz testing agent 40 via the input subsystem 30, the message containing the address of the location for placing the hook. In some examples, the network manager 42 of the fuzz testing agent 40 receives the message, and then the fuzz testing agent 40 parses the received message to find the address offset of the code block and the address offset of the point of interest. In some examples, the fuzz testing agent 40 then adds a base address (such as an ASLR base address) to the address offset to identify the actual memory address of the code block and the actual memory address of the point of interest.
[0151] In some examples, the fuzz testing agent 40 changes the access permission of the text section of the binary executable file to "write". In some examples where the DUT 115 is a Linux system, the changing of the access permission is performed using the Mprotect application programming interface (API). In another example where the DUT 115 is an embedded system, the changing of the access permission is performed using the memory protection module API.
[0152] In some examples, as described above, the fuzz testing agent 40 adds a hook by replacing the opcode at the corresponding address with a branch command to the event handler 41. In some examples where multiple event handlers 41 are provided, each event handler 41 is associated with a corresponding event type from among a plurality of event types. For example, as described above, the event types may include POI events, coverage events, CFI events, and statistical events. In such an example, the fuzz testing agent 40 includes four event handlers 41, namely a first event handler 41 associated with a POI event, a second event handler 41 associated with a coverage event, a third event handler 41 associated with a CFI event, and a fourth event handler 41 associated with a statistical event.
[0153] Similarly, each hook branches to a corresponding event handler 41 according to the type of hook. For example: a POI event hook is arranged at a point of interest (e.g., Figure 4F H1 in ), and thus branches to the event handler 41 associated with the POI event; the overlay event hook is arranged at the beginning of the code block (e.g., Figure 4F H5, H3 and H2 in ), and thus branches to the event handler 41 associated with the coverage event; and the CFI event hooks are arranged at points where there may be control flow or security errors (e.g., Figure 4F Hook H4 in the middle).
[0154] In some examples, the CFI event hook is added after the POI event hook is reached. In particular, in such examples, the corresponding event handler 41 receives an indication from the POI event hook that the corresponding point of interest has been reached. In response to receiving such an indication, the fuzz testing agent 40 adds the CFI event hook to the corresponding portion of the code. In some examples, the fuzz testing agent 40 removes the POI event hook that was reached and replaces it with a CFI event at the same location. Therefore, in such examples, the POI event hook is used to identify when the process flow reaches a point of interest, and the CFI event hook is used to perform actual fuzz testing on the corresponding point of interest to detect the CFI event.
[0155] In some examples, each hooked branch instruction includes a branch-with-link instruction. As known to those skilled in the art, the branch-with-link instruction branches to a predetermined address while saving a return address. In some examples, the return address of each hook is stored together with the corresponding opcode replaced by the hook, so that the fuzz testing agent 40 can remove the corresponding hook and return the replaced opcode to its original address.
[0156] In some examples, the opcode replaced by the corresponding hook is stored within the corresponding event handler 41. In such examples, upon reaching the corresponding hook, the process branches to the corresponding event handler 41, and then the corresponding event handler 41 identifies the location of the corresponding hook. In some examples, the hook is identified by comparing the return address received from the branch instruction with the table containing the return address of the replaced opcode. In such examples, the replaced opcode is then executed inside the corresponding event handler 41. For example, for an opcode that includes a comparison of the value of a register with a predetermined value, the corresponding event handler 41 performs the corresponding comparison and then returns to the appropriate return address. Advantageously, running the replaced opcode inside the corresponding event handler 41 is faster than storing the replaced opcode in a different location, finding the location, and branching to the location to execute the opcode.
[0157] In some examples, each event handler 41 is a function, and at the end of the execution of the function, it returns to the caller. In such examples, the corresponding event handler adjusts the return address so that it continues to the next opcode, that is, the amount of the return address offset is the number of bytes between each opcode. For example, in an ARM32 environment, in the case where the return address is 0x100, the return address will be adjusted to 0x104.
[0158] In another example, the replaced opcode is stored in a different memory address, and the corresponding event handler 41 branches to the appropriate address to reach the replaced opcode.
[0159] As described above, in some examples, when a certain type of hook is reached (such as an override event type hook), the corresponding event handler 41 removes the hook and puts the replaced opcode back in its original location.
[0160] In step 610, for each point of interest, a fuzzy test is performed on the particular point of interest within a predetermined test time. In some examples, the time it takes to reach the point of interest (which may take time to reach the point of interest if there are conditional checks along the way) is included in the maximum allowed test time. In another example, the predetermined test time is defined as the maximum allowed time for attempting to reach the point of interest.
[0161] As described above, fuzz testing unit 51 controls fuzzer data generator 20 to provide data units to input subsystem 30. In some examples, the fuzz testing unit modifies data units for fuzz testing according to a genetic algorithm or other suitable fuzz testing algorithm as known to those skilled in the art.
[0162] In some examples, after inserting a data unit through the input subsystem 30, the fuzzer evaluation function block 50 has the following possible situations for receiving information from the network manager 42 of the fuzz testing agent 40: A. No information is received, that is, no hook is reached; B. Information indicating a POI event; C. Information indicating a coverage event; or D. Information indicating a CFI event. For each data unit sent, the network manager 42 can receive information about multiple hooks that are reached. In some examples, the control unit 52 stores information about the initiated events in a buffer, and after a predetermined test time, or after a predetermined number of hooks have been reached, the information in the buffer is stored in the memory 140.
[0163] In some examples, each hook has a corresponding score associated with a corresponding point of interest. In some examples, the coverage event hook has a score associated with the distance from the point of interest. For example, for Figure 4F , hook H5 (which is a coverage event hook) has a score of 1 associated with the point of interest FUNC2 because it is located in the first code block in the call tree of FUNC2. Similarly, hook H3 (which is a coverage event hook) has a score of 2 because it is located in the second code block in the call tree of FUNC2. Similarly, hook H2 (which is a coverage event hook) has a score of 3 because it is located in the third code block in the call tree of FUNC2. Although described above in an example where the closer the coverage event hook is to the point of interest, the higher its score, this is not meant to be limiting in any way. In another example, a POI event hook (such as hook H1) has a higher score than a coverage event hook, and a CFI event hook (such as hook H4) has a higher score than a POI event hook. The following table shows Figure 4F An example of an event hook:
[0164] Table 1
[0165]
[0166]
[0167] Wherein, as described above, the timestamp indicates a timestamp generated when the process reaches the corresponding hook, and the address indicates the address of the hook. As will be described below, the score is used by the fuzzer evaluation function block 50 to generate a coverage report and / or to adjust the fuzz test of the interest point.
[0168] In some examples, to identify how much coverage value has been achieved, a total coverage score is defined as a predetermined function of different coverage event hooks that have been reached, where coverage event hooks with different scores exhibit different weights. In some examples, a total coverage score is determined for each data unit. In another example, a total coverage score is determined at the end of a fuzz testing session to determine the coverage value achieved.
[0169] In some examples, the coverage score is determined as follows:
[0170] A. Arrival hooks with level 1 hooks are defined as having a predetermined score. Level 1 hooks are defined as covered event hooks that are farther from the point of interest ( Figure 4F The score of the level 1 hook is represented as "score_level_1_hook".
[0171] B. “Score_level_2_hook” is defined as: (number of level 1 hooks)*score_level_1_hook+1.
[0172] C. "Score_level_3_hook" is defined as: (number of level 1 hooks)*score_level_1_hook+(number of level 2 hooks)*score_level_2_hook+1. Level 2 hooks are defined as the coverage event hook in the second code block in the call tree of the point of interest ( Figure 4F Hook H3 in the middle).
[0173] D. Further define the scores for each level based on the above.
[0174] Therefore, each event has its own score, and the data unit can be adjusted according to the score of each event to reach the corresponding interest point.
[0175] According to the received information about the arrived hook, the fuzz test unit 51 accordingly adjusts the data unit of the fuzzer data generator 20. For example, for each POI, in step 620, the control unit 52 of the fuzzer evaluation functional block 50 determines whether the corresponding point of interest has been reached, that is, whether a POI event associated with the corresponding point of interest has been initiated.
[0176] If the control unit 52 determines that the corresponding data unit does not reach the corresponding point of interest, then in step 630, a function-level fuzzy test is performed on the code block closest to the corresponding point of interest. Figure 4F If the function level fuzz test is performed on the block 0x122, the function level fuzz test is performed on the block 0x122. In some examples, the closest code block is identified according to the score of the coverage hook at the beginning of the corresponding code block. For example, the coverage event hook exhibiting the highest score (or the second highest score) will be located in the code block immediately before the code block containing the point of interest.
[0177] In some examples, function-level fuzz testing performed by the fuzz testing agent 40 creates a snapshot of the target CPU internal state (registers and memory) and sends the snapshot to the fuzzer evaluation function block 50. In some examples, in the case of hardware-related functions (e.g., ECU peripherals), the fuzz testing agent 40 sends the relevant peripheral information for function-level fuzz testing to use to simulate the hardware-related functions.
[0178] In some examples, the control unit 52 of the fuzzer evaluation function block 50 sends the snapshot information and optionally other additional information to the network manager 42' of the fuzz testing agent 40' running in the emulator 115'. In some examples, the additional information includes any of the following: the address of the point of interest; the number of pointer bytes copied; or whether a CFI event is detected.
[0179] In some examples, in the presence of a function being fuzzed and a hardware dependency that is not simulated, the control unit 52' does not crash or stop function-level fuzz testing due to a lack of hardware dependencies, but instead requests from the fuzzer evaluation functional block 50 to perform network-level fuzz testing on the DUT 115 until it reaches the function that calls the hardware dependency, and then the fuzz test agent 40 sends the hardware dependency information to the fuzzer evaluation functional block 50. The control unit 52 of the fuzzer evaluation functional block 50 then forwards the information to the control unit 52' in the simulator 115'. The control unit 52' then updates the fuzz test agent 40' to simulate the hardware-related functions, and when the hardware dependency is called, the fuzz test agent 40' returns the hardware dependency value (received from the DUT 115) to the function.
[0180] In one embodiment, function-level fuzz testing is performed using common utilities for function-level fuzz testing (such as AFL or libfuzzer). In some examples, the fuzz testing agent 40' wraps the function under test (FUT) and monitors its state (runtime duration, return value, memory, etc.).
[0181] In some examples, the control unit 52' controls the fuzzer unit 51' to input a value into the corresponding code block in order to reach the point of interest. If the code block includes one or more conditional checks, the value is input until the correct value for overcoming the conditional check (or multiple conditional checks) is found. Therefore, the control unit 52' and the fuzzer unit 51' continue to perform function-level fuzz testing until the point of interest is reached.
[0182] After completing the function-level fuzz testing, the network manager 42' sends the values used to reach the point of interest to the fuzzer evaluation function block 50. The fuzzer evaluation function block 50 then uses these values to control the fuzzer data generator 20 to generate data units containing these values. In particular, in some examples, the data units are repeatedly updated and sent until the determined argument values of the corresponding function are achieved. In some examples, the fuzzer evaluation function block includes a predetermined algorithm for updating the data units in response to changes in the function parameters so that the difference between the function parameters and the determined argument values remains small.
[0183] In step 640 , the fuzzer evaluation function 50 then checks again whether a point of interest has been reached.
[0184] If a point of interest is reached in step 630 or step 610, function-level fuzz testing is performed in step 650 to identify CFI events. Advantageously, function-level fuzz testing is performed faster than network-level fuzz testing. Therefore, identifying CFI events in function-level fuzz testing will be faster than identifying CFI events in network-level fuzz testing. In addition, while function-level fuzz testing is performed to identify CFI events, network-level fuzz testing can be continued to identify other POI events.
[0185] During function-level fuzz testing, the control unit 52' and the fuzz testing unit 51' perform fuzz testing on points of interest (e.g., functions) using varying function parameters to identify abnormal events such as memory corruption, run duration greater than a predetermined time threshold, attempts to access disallowed memory (e.g., segfaults), and the like.
[0186] In some examples, the event handler 41' stores the function parameters that caused the event in a dedicated buffer. In some examples, the function parameters are stored together with the identifiers of their corresponding registers. Since the parameters of the function can be pointers, in some examples, the event handler 41' verifies that each parameter value is a legal address in the memory space. If the process memory has such a value as an address, the event handler 41' copies the corresponding number of bytes from the address to the buffer. In some examples, the corresponding number of bytes is a predefined predetermined number.
[0187] In some examples, function parameters are stored in memory or sent by the network manager 42' to the fuzzer evaluation function block 50. In some examples, the decision whether to store event information or send event information is based on configuration information received at the beginning of function-level fuzz testing. In some examples, function-level fuzz testing of points of interest is run until a predetermined test time has passed. If function parameters are present at each CFI event, then the fuzzer evaluation function block 50 now contains register values that can be used to cause a CFI event at the point of interest.
[0188] In step 660, the fuzzer evaluation function block 50 determines whether a CFI event occurred during the function-level fuzz testing. If at least one CFI event occurs, the probability of the CFI event actually occurring is checked in steps 670 and 680, respectively, as described above in conjunction with steps 430 and 440. In other words, the CFI event is verified to determine whether it is a real CFI event or merely theoretical. In particular, step 430 corresponds to step 670, and step 440 corresponds to step 680. Although both steps 670 and 680 are described as being executed, this is not meant to be limiting in any way. In another example, only one of steps 670 or 680 is executed. In another example, each point of interest has defined for this which of steps 670 or 680 should be executed, or whether both should be executed. In another example, for one or more points of interest, neither step 670 nor step 680 is executed.
[0189] In step 670, network-level fuzz testing is used to check the probability of a CFI event occurring. Specifically, in some examples, the fuzzer evaluation function block 50 has previously received function parameters that cause CFI events, as described above. These function parameters are used as target values. In some examples, the fuzzer evaluation function block 50 instructs the fuzz testing agent 40 to perform a network-level fuzz test on the code block containing the point of interest ( Figure 4F An information leakage event hook is added at the beginning of a block (block _0X111 in ). As used herein, the term "information leakage event hook" refers to a hook that copies the parameter value of a function from a corresponding register or memory address. For example, in the ARM32 instruction set, function parameter values are typically stored in registers r0, r1, r2, etc. In some examples, arranging the information leakage event hook at the beginning of a code block can provide higher resolution capabilities because a function can include multiple code blocks. However, this is not meant to be limited in any way. In some examples, one or more information leakage event hooks are arranged at the beginning of the corresponding function.
[0190] In some examples, the fuzzer evaluation function 50 starts the network-level fuzz testing by instructing the fuzzer data generator 20 to start a fuzz testing session using the data units that reach the point of interest in step 620 (or step 640).
[0191] In response to reaching the corresponding information leakage event hook, in some examples, the hook branches to the corresponding event handler 41 associated with the information leakage event hook. In some examples, the corresponding event handler 41 updates the event buffer with the current function parameter value. In some examples, the fuzz testing agent 40 sends the event data received from the information leakage event hook to the fuzzer evaluation function block 50.
[0192] In some examples, the fuzzer evaluation function block 50 uses the event information as a score value for an optimization algorithm for updating the data unit. In some examples, the optimization algorithm includes a genetic algorithm, such as an adaptive heuristic search algorithm. In another example, other optimization algorithms can be used, such as the algorithm provided by libfuzzer, which is commercially available from Google LLC of Mountain View, California, USA.
[0193] In some examples, a distance value is defined by comparing a current parameter value with a target parameter value received from the simulator 115'. The distance value acts as a score for the data unit. For each data unit that is sent and arrives at the point of interest, the current parameter value is compared with the target parameter value, and the distance value between them is defined as the score for the corresponding data unit. The optimization algorithm (in the fuzzy testing unit 51) uses this feedback mechanism and the score to find one or more data units that can result in a parameter value equal to the target parameter value.
[0194] If the predetermined test time has passed and no such data unit is found, the data unit with the lowest score (i.e., the lowest distance value) is reported by the fuzzer evaluation function block 50 in step 690. If the fuzz testing session time has passed and no packet is found, the data unit with the highest score is reported / stored in step 690. If such a data unit is found, the corresponding data unit is reported / stored in step 690. In some examples, the control unit 52 of the fuzzer evaluation function block 50 stores all data units and their scores in the memory 140.
[0195] In step 680, function-level fuzz testing is used to check the probability of a CFI event occurring. In some examples, the fuzzer evaluation function block 50 instructs the fuzz testing agent 40 to add a coverage event hook to the beginning of a code block (the code block calls the code block including the point of interest), as described above in conjunction with step 630. In some examples, as described above, the instruction from the fuzzer evaluation function block 50 to the fuzz testing agent 40 is sent via a packet targeting the IP and port of the fuzz testing agent 40, and the payload of the packet includes the address where the hook should be placed and the type of hook to be placed.
[0196] When the process flow reaches a new hook, in some examples, the fuzz testing agent creates a snapshot of the memory space and sends the snapshot to the fuzzer evaluation function block 50, as described above. As described above, the fuzzer evaluation function block 50 sends the snapshot information and optionally other additional information to the network manager 42 running in the emulator 115';.
[0197] Then, the fuzz testing unit 51 'performs function-level fuzz testing (as described above) to try to find function parameters that are found to cause CFI events. In particular, the goal of this stage is to find the case where the previous code block calls the POI block (i.e., the code block containing the point of interest) with the same parameters that caused the CFI event.
[0198] In some examples, the fuzz testing unit 51 'tests different sets of parameters for the previous code block and uses the parameter values from the CFI event as targets. In some examples, the difference between the current value sent to the POI block and the value found in the CFI event is defined as a corresponding score. In some examples, the fuzz testing unit 51 ' applies an algorithm that aims to maximize the score by changing the value.
[0199] If the predetermined time period has passed, but no such parameter value for calling the POI block is found, the control unit 52' updates the fuzzer evaluation function block 50 that no parameter was found. If such a parameter value is found, the control unit 52' updates the fuzzer evaluation function block 50 with the identified parameter value. In some examples, the network-level fuzz test of step 670 is then performed as described above based on the identified parameter value.
[0200] In another example, the function-level fuzz test of step 680 is performed again for the code block before the code block that has just been fuzz tested, so as to find the parameter value of calling the corresponding code block while maintaining the corresponding parameter value that caused the CFI event. As described above, in some examples, the fuzz test agent 40 adds a hook to the previous code block and performs function-level fuzz testing based on a snapshot obtained when the new hook of the previous block is reached. Therefore, in some examples, function-level fuzz testing is repeatedly performed, going through consecutive code blocks in reverse, until the first code block of the binary executable file 110 is reached or until the function-level fuzz test can no longer reach another code block. In such an example, the corresponding parameter value and the corresponding block arrival are reported and / or stored in step 690.
[0201] Figure 5A A high-level block diagram of an example of a system 700 for fuzz testing is shown. System 700 is similar to system 300 in all respects, except that fuzzer data generator 20 and fuzz test unit 51 are internal to DUT 115, while control unit 52 is external to DUT 115. In such an example, input subsystem 30 is not required because data units are provided from fuzzer data generator 20 to binary executable file 110 via local host interface 710. Advantageously, this reduces the network latency that exists when data units enter through input subsystem 30.
[0202] In some examples, the fuzzer data generator 20 sends the data unit to the binary executable file 110 via the local host interface 710 implemented as a loopback network interface. In some examples, the fuzz test agent 40 adds a hook to the initialization function for configuring the communication between the fuzzer data generator 20 and the binary executable file 110. In another example, the fuzz test agent 40 adds a hook at socket.bind. As described above, in such an example, socket.bind is replaced with a branch instruction to the corresponding event handler 41, and socket.bind runs within the corresponding event handler 41. The event handler 41 changes socket.bind to change the source monitored through the loopback network interface. In some examples, the event handler 41 changes socket.bind to change the allowed monitoring source to "0.0.0.0", that is, all monitoring sources are allowed.
[0203] In another example, a kernel module including a Linux net-filter is used to modify the data unit generated by the fuzzer data generator 20. As used herein, the term "kernel module" refers to an object file containing code that can extend kernel functionality at runtime, as known to those skilled in the art. In such an example, the generated data unit is received via a loopback network interface and sent to the kernel module via a network filter (netfilter) input chain. The net-filter then appropriately modifies the IP address of the data unit and returns the data unit to the network filter input chain (as known to those skilled in the art), and the modified data unit is then sent to the binary executable file 110.
[0204] In some examples, DUT 115 can be replaced with a virtual machine (VM) or a virtual container (such as a Docker container commercially available from Docker Inc. of Palo Alto, California, USA). Figure 5B As shown, a system 800 for fuzz testing is provided. The system 800 is similar to the system 700 in all aspects, except that the DUT 115 is replaced with a plurality of virtual environments 810 (such as VMs or virtual containers).
[0205] In some examples, the control unit 52 receives an instrumented binary executable file and multiple configuration files or messages. In particular, each configuration file / message indicates which parts of the binary executable file are to be fuzz tested and which parameters are used for fuzz testing, as described above (e.g., the number of data units, the time used for fuzz testing, TARA information, etc.). Therefore, the control unit 52 fuzz tests each section of the binary executable file in a separate virtual environment 810. In some examples, each virtual environment 810 can be accessed through a local network interface. In another example, each virtual environment 810 is accessed through a network via a corresponding IP address. In some examples, where a local network interface is provided, each virtual environment 810 has a dedicated fuzz testing unit, as described above in conjunction with the fuzz testing unit 51 of the system 700.
[0206] 6A to 6F Various high-level block diagrams of examples of a proxy service-based fuzz testing system are shown. In some examples, the steps of the proxy service-based fuzz testing system include: a binary file analysis phase; a fuzzer generation phase; and a runtime phase.
[0207] In some examples, a configuration file is created that includes information about the address of each logical block in the binary file. Optionally, the configuration file includes a list of corresponding addresses.
[0208] In some examples, for each block that has a condition within the block (as described above), an offset to the address of the condition and a number of parameters checked in the condition are added to the configuration file.
[0209] In some examples, the configuration file also includes a list of all entry points to the binary file. The entry points may include calls to read functions, receive functions (e.g., recvfrom), and other similar entry points.
[0210] In some examples, one or more hooks are placed at various points of interest in a binary executable file using a configuration file, as described above. In some examples, hooks are added only at some points of interest, as described above. In some examples, a list of points of interest to add hooks to is saved in a configuration file. In some examples, points of interest include, but are not limited to, entry points of processes, entry points of blocks, and / or conditional checks, as described above.
[0211] In some examples, the entry of each block is replaced with a hook according to the information of the configuration file, as described above. In some examples, each hook arranged at the entry of each block includes a call to / branch to the corresponding code, which sends a coverage event message to the proxy service module 820, such as Fig. 6A As used herein, the term "overwrite event message" refers to information about an overwrite event, as described above. The overwrite event message indicates arrival at a corresponding hook. In some examples, the corresponding overwrite event message associated with each hook includes an identifier of the corresponding hook / block. As described above, the process of the binary executable file continues.
[0212] In some examples, each conditional opcode is replaced with a corresponding hook according to the information of the configuration file. As used herein, the term "conditional opcode" refers to an opcode with a conditional check, as described above. In some examples, each hook that replaces a conditional opcode includes a call to the corresponding code / a branch to the code, which sends a conditional event message to the proxy service module 820. As used herein, the term "conditional event message" refers to information about conditional events, as described above. In some examples, the conditional event message includes the corresponding register value of the condition (e.g., the corresponding variable value and parameter value of the condition). In some examples, as described above, the code also executes the condition. As described above, the process of the binary executable file continues.
[0213] In some examples, based on information from the configuration file, a hook is added for one or more calls to the entry point. In some examples, the hook includes a branch to a corresponding code that receives data from a communication channel. In particular, as will be described below, the communication channel is opened to receive the data unit.
[0214] An illustrative example of a configuration file could be as follows:
[0215] Table 2
[0216] ID Offset type parameter Previous block 1 0x100 cover none none 2001 0x108 condition R0,R1 0x100 2002 0x132 condition R0,#223 0x100 2 0x148 cover none 0x100
[0217] In some examples, a CFI monitor is added to detect memory corruption and other CFI events. If a CFI event occurs, a CFI event message is sent to the proxy service module 820 as described above. As used herein, the term "CFI event message" refers to information about a CFI event that occurred, optionally including details of the event.
[0218] In some examples, the event handler is embedded in a binary executable file, as described above. In some examples, the event handler sends an event message to a communication channel.
[0219] In some examples, as will be described below, based on information of the configuration file, code for the proxy service module 820 is generated. In some examples, a first portion of the code for the proxy service module 820 is independent of the configuration file information, and a second portion of the code for the proxy service module 820 is dependent on the configuration file information. Thus, the proxy service module 820 can be pre-programmed and then updated in response to a received configuration file.
[0220] In some examples, a fuzzer grammar and a fuzzer seed are generated in response to received user input.
[0221] In some examples, as described above, the binary executable file 110 runs in an execution context 825, which is, for example, a DUT or a virtual environment, such as a virtual machine or an emulator. In some examples, as will be described below, the binary executable file 110 receives data from a communication channel. The binary executable file 110 then processes the incoming data.
[0222] When the logic flow control reaches the hook at the beginning of the block, an overlay event message (as described above) is generated and sent to the proxy service module via the communication channel. Similarly, when the logic flow control reaches a condition, a condition event message (as described above) is generated and sent to the proxy service module via the communication channel.
[0223] The proxy service module 820 includes source code so that it can be compiled to support various fuzzers, including coverage-guided fuzzers such as AFL as known to those skilled in the art (e.g. Figure 6BIn some examples, as will be described below, the proxy service module 820 communicates with the instrumented binary executable file 110 using a communication channel.
[0224] The proxy service module 820 receives event messages (e.g., overlay event messages and conditional event messages) from the instrumented binary executable file 110. In some examples, the event handler of the proxy service module includes being configured to wait for a predetermined period of time (preferably measured in microseconds) after receiving an event to determine that it has received the last event for the data unit sent, and only after the timeout, it will send the next data unit of the fuzz testing process. In some examples, the waiting is performed in the following situations, but not limited to these situations: when there are dependencies between events; when the server uses a request-response technique to send data units; and / or the binary executable file includes multiple threads, and the transmitted data unit can trigger events from more than one thread.
[0225] As will be described below, the proxy service module 820 provides input to the fuzzer 830 (e.g., AFL, Libfuzzer, or AFL++) in response to received events. In some examples, the proxy service module 820 includes multiple branch instructions. The proxy service module is described herein as including multiple functions, each of which is called in response to a corresponding event message, however, this is not meant to be limiting in any way. In some examples, the proxy service module 820 includes multiple conditions (such as "if" statements), each of which is branched in response to a corresponding event message. In some examples, a counter can be incremented or other suitable actions can be taken by a condition. In some examples, the proxy service module includes a lookup table that calls a function when a corresponding event message is obtained.
[0226] In some examples, the proxy service module includes an array of all functions, such as the following array: void(*Funcs[NUMBER_OF_EVENTS])()=
[0227] {function_handler_100,
[0228] function_handler_148,…};
[0229] In some examples, for each event listed in the configuration file, a corresponding function is generated. In some examples, hooks are placed at each block entry point and each conditional check in the binary executable file, however, the configuration file contains a dedicated list of partial events to be used for fuzz testing. In such an example, the event handler of the proxy service module 820 will ignore other events, thereby focusing the fuzz testing on the flow of one or more predetermined points of interest.
[0230] In some examples, when the proxy service module 820 is started, the initialization step includes reading a configuration file from memory. In some examples, the event handler of the proxy service module 820 will read a list of event IDs from the configuration file and only take action on received events whose IDs are in the list. In some examples, if the configuration file does not include a list of events, or such a list is empty, the event handler of the proxy service module 820 will take action on each event. In some examples, this provides the ability to change the points of interest being fuzzed by simply creating a new configuration file.
[0231] The fuzzer 830 is designed to update and output data units so as to reach as many functions as possible, as known to those skilled in the art of coverage-guided fuzz testing. Using the proxy service module 820, the fuzzer 830 attempts to improve the coverage value within the proxy service module 820, i.e., the number of functions reached, each of which is called in response to a corresponding event within the binary executable file 110. In some examples, this allows the binary executable file 110 to be indirectly fuzzed using a standard coverage-guided fuzzer even if the binary executable file cannot be fuzzed directly due to certain constraints (e.g., lack of source code).
[0232] In some examples, for each condition event message, the corresponding function includes a condition check and a call to a pair of dedicated functions. In an illustrative example, as described in Table 2 above, a condition event may have an offset of 0x132, and the condition event message includes a parameter value R0 and a hard condition value, which is equal to 223. In such an illustrative example, the corresponding function may look like this:
[0233] void function_handler_100(void){
[0234] if(R0==223){
[0235] success_132();
[0236] }else{
[0237] Failure_132();
[0238] }
[0239] Void function_success_132(void){}
[0240] Void function_Failure_132(void){}
[0241] In such an example, only when the original condition has been met, the corresponding success function will be called, otherwise the failure function will be called. The fuzzer 830 is configured to continue adjusting the data unit until the success function is called.
[0242] By checking the original condition, the fuzzer 830 can continue fuzz testing until the condition is reached. In some examples, the fuzzer includes a dedicated algorithm that continuously updates the data unit to minimize the distance between R0 and 223. Although described above in conjunction with specific numerical examples, this is true for all condition checks. In addition, instead of a fixed number (e.g., 223), the conditional event message may include a non-fixed value, such as R1. In such an example, fuzz testing continues until the value of R0 is equal to the value of R1. Therefore, regardless of the conditions, a coverage-guided fuzzer can be used to fuzz test binary executable files.
[0243] In one illustrative example, when a conditional event message is received, the code of the proxy service module may look like this:
[0244] Event_handler(){
[0245] Int E=Receive_event_from_communication();
[0246] If(E is condition_event){
[0247] Set R0 and R1 according to the content received from the event message
[0248] }
[0249] Funcs[E]();
[0250] }
[0251] In such an illustrative example, these values are set to register values associated with conditional checks in the binary executable file, and then the corresponding function is called.
[0252] In some examples, in the event of a crash, the tested binary executable 110 will no longer run. This may mean that the crash is not reported to the proxy service module 820. Therefore, in some examples, the system also includes a monitor that detects runtime failures of the tested binary executable and reports the runtime failures to the proxy service module 820 using a failure event message. In response to receiving the failure event message, in some examples, the proxy service module 820 calls an error function that crashes the proxy service module 820, whereby the fuzzer 830 sees the crash.
[0253] In some examples, such a monitor includes a debugger, which will also allow postmortem analysis.
[0254] In some examples, the monitor can perform any one or a combination of the following functions: reporting a crash, including the cause of the crash, such as a seg fault; providing a core dump in response to a crash (for performing post-mortem analysis); and injecting trace points, such as for counting the size of allocated memory and the number of free calls, to detect memory leaks.
[0255] In some examples, such as Figure 6C As shown, where the proxy service module 820 and the binary executable file 110 are running in the same execution context 825, the shared memory can be used to forward the data unit to the binary executable file under test and report the event back to the proxy service module. The data unit is referred to as a packet hereinafter, however, this is not meant to be limiting in any way, and any type of data transmission can be used without exceeding the scope of this disclosure.
[0256] In some examples, two separate queues are used: a packet queue; and an event queue. In some examples, the packet queue buffers packets provided by the proxy service module 820. In some examples, the packet injection engine pops the packet from the queue and injects it into the receiving mechanism of the binary executable file 110 under test, such as by linking against a prepared recv call.
[0257] In some examples, an event handler embedded in the binary executable file 110 collects events (as described above) and pushes these events to an event queue. The proxy service module 820 can then pop the events as needed.
[0258] In some examples, such as FIG. 6D to FIG. 6EAs shown, the proxy service module 820 and the binary executable file 110 under test do not run in the same execution context, and the communication socket is used to forward packet data to the binary executable file 110 under test and report events back to the proxy service module 820. In some examples, such as Fig.6D As shown, the proxy service module 820 runs on the same machine as the binary executable file 110. Alternatively, Fig. 6E As shown, the proxy service module 820 runs in one machine, while the target binary executable file 110 runs in a different machine.
[0259] In some examples, two separate queues are used: a packet queue 840; and an event queue 850. In some examples, the packet queue 840 buffers packets provided by the proxy service module 820. In some examples, the packet injection engine pops the packet from the queue 840 and sends it to the binary executable file under test 110 using a socket. In the case where the binary executable file 110 does not use sockets for communication, in some examples, a socket listener (for listening to network communications) is added to the binary executable file under test 110. In some examples, the socket listener (also called a "network client") receives the packet and injects it into the entry point of the binary executable file. For example, it can provide data for specific code added to the entry point.
[0260] In some examples, an event handler embedded in the binary executable collects events (as described above) and sends them to the proxy service module 820 using a dedicated UDP message. In some examples where the proxy service module 820 and the binary executable 110 are running on the same machine, the UDP message can be a simple UDP message to "localhost". In some examples where the proxy service module 820 and the binary executable 110 are running on separate machines, the proxy service module 820 sends the message to a remote IP address. In some examples, the proxy service module 820 implements a UDP listener to receive UDP packets.
[0261] In some examples, such as Fig. 6F As shown, when the firmware or portable operating system interface (POSIX) binary executable file 110 is fuzz tested on its native target hardware, the events are read by the debugger 860. In some examples, the packets are forwarded to the network adapter of the target hardware.
[0262] In some examples, two separate communication channels are used: a communication channel for packets; and a communication channel for events. In some examples, the packet queue buffers packets provided by the proxy service module 820, as described above. In some examples, the network module pops the packet from the queue 840 and forwards it to the network adapter of the target hardware. In some examples, forwarding the popped packet is accomplished while maintaining the rate limit of the network adapter.
[0263] In some examples, the instrumented binary executable records the event into a global buffer. In some examples, the global buffer is periodically polled using a debugger 860. In some examples, the software controlling the debugger 860 forwards the event to an event server in the execution context of the proxy service module. In some examples, the event is then sent to a queue manager, as described above.
[0264] Figure 7 A high-level flow chart of a signal-based fuzz testing method is shown. As used herein, the term "signal-based fuzz testing" refers to fuzz testing a target based on changes made to a signal. In particular, each signal has its own predetermined position within a corresponding payload. Therefore, fuzz testing is performed by changing (e.g., by mutation) bits in a corresponding position of the payload, and the corresponding position represents the location of the corresponding signal that is typically sent to a binary executable file. It is envisioned that different signals can be associated with different origins and / or destinations, and therefore, in some examples, each signal within a data unit (or network packet) is defined based on a position within the data unit and one or more identifiers of the data unit. As described above, fuzz testing includes continuously adjusting the data unit, and then inputting the data unit into the target.
[0265] In some examples, signal-based fuzz testing is performed on one or more predetermined signals in stage 900. In some examples, the signal-based fuzz testing is performed as described above in conjunction with any of systems 10, 215, 300, 700, or 800. In some examples, the signal-based fuzz testing is performed using a different fuzzifier, such as an AFL fuzzifier.
[0266] In some examples, in stage 910, when a hook is reached, the corresponding hook outputs information associated with the corresponding point of interest, as described above. In some examples, the fuzzer evaluation function block 50 stores and / or outputs information about the signal and the one or more hooks reached. In some examples, for each signal being fuzz tested, the fuzzer evaluation function block 50 outputs a list of hooks and / or points of interest reached.
[0267] In some examples, for each signal being fuzz tested, the fuzzer evaluation function block 50 determines whether the corresponding signal reaches one or more hooks in the hook subset, and in some examples, the fuzzer evaluation function block 50 also outputs an indication of whether the one or more hooks are reached. In some examples, the hook subset is associated with high-risk points of interest. Therefore, it is determined whether the corresponding signal reaches any such high-risk points of interest.
[0268] In one illustrative example, the output of the fuzzer evaluation function block 50 may include the following fields:
[0269] Table 3
[0270] Binary file ID Signal ID List of arrived hooks B1 S1 H1, H3, H7... B1 S2 H2, H3, H9... B1 S3 H4, H5, H11...
[0271] As shown, in such an example, a list of hooks in each binary executable file that are reached by each signal may be provided.
[0272] In some examples, the output of the fuzzer evaluation function block 50 (such as the output described in Table 3) is output to an external system, an external network, and / or a user terminal.
[0273] In some examples, in stage 920, one or more signals of stage 900 include multiple signals, i.e., a signal group, each of which is located at a different position of the same data unit / payload. In some examples, signal-based fuzz testing is performed together for the signal group. In some examples, this includes changing the bits of all signals into a single data block. In some examples, this includes individually changing the bits of one or more signals in the multiple signals according to a predetermined rule. It should be noted that certain values of a signal may only reach specific points of interest if other signals have one or more specific values. Therefore, fuzz testing the signals together (as a single data block, or in a predetermined order) can help reach the corresponding points of interest. In some examples, the fuzzifier evaluation function block 50 outputs information about the hooks (and / or points of interest) reached by the signal group together.
[0274] In some examples, in stage 930, based at least in part on a determination that the corresponding signal reaches one or more specific points of interest (such as high-risk points of interest), the fuzzifier evaluation function block 50 determines that further fuzz testing should be performed on the corresponding signal. In some examples, the points of interest are defined as high risk by the fuzzifier evaluation function block 50 and / or external input. In some examples, the fuzzifier evaluation function block 50 defines the points of interest as high risk based at least in part on data received from the outside. In some examples, the fuzzifier evaluation function block 50 receives TARA information, and the definition of the points of interest as high risk is based at least in part on the received TARA information. In some examples, each point of interest is assigned a corresponding risk value (assigned by the external input and / or by the fuzzifier evaluation function block 50), and a threshold is defined such that each point of interest assigned a risk value greater than the threshold is defined as a high-risk point of interest.
[0275] A high risk point of interest may be any point of interest that is defined as high risk, including but not limited to: access points; access points to software / hardware with a high risk value, optionally determined by a risk assessment (such as TARA); and / or points of interest with known vulnerabilities, such as having a known Common Vulnerabilities and Exposures (CVE) identifier.
[0276] In some examples, as long as a particular point of interest has not been reached, only a predetermined maximum amount of changes are made to the signal for fuzz testing. However, once a particular point of interest has been reached, in some examples, a greater amount of changes may be made to the signal for fuzz testing. Thus, intelligent fuzz testing is provided, wherein if the signal reaches a predetermined point of interest, the signal is fuzzified to a greater extent, and if the signal does not reach the predetermined point of interest, the signal is fuzzified to a lesser extent.
[0277] In some examples, the further fuzz testing includes fuzz testing the signal to reach additional points of interest, which are optionally points of interest accessed through the first point of interest. For example, the particular point of interest can be an access point to the corresponding system, such as an access point to a modem. Once the access point is reached, further fuzz testing is performed on the corresponding signal to reach additional points of interest within the accessed system. In some examples, the further fuzz testing includes fuzz testing the signal to generate: an error or failure in the system; and / or a heavy CPU load. In some examples, the further fuzz testing includes fuzz testing the signal for at least a predetermined period of time.
[0278] Figure 8A high-level flow chart of a method for identifying statistical independence of multiple signals is shown. The following description will be combined with an example related to analyzing the statistical independence of two signals, however, this is not meant to be limiting in any way, and any number of signals can be used to determine the statistical independence of any number of signals without exceeding the scope of the present disclosure.
[0279] In some examples, at stage 1000, signal-based fuzz testing is performed on the first signal, as described above. In some examples, the signal-based fuzz testing is performed as described above in conjunction with any of systems 10, 215, 300, 700, or 800. In some examples, the signal-based fuzz testing is performed using a different fuzzer, such as an AFL fuzzer. In some examples, when the first signal is modified, the second signal is not modified.
[0280] In some examples, as described above, the fuzzer evaluation function 50 determines which hooks the data unit has reached. In some examples, the fuzzer evaluation function 50 determines other effects of the first signal, such as a high load on the CPU.
[0281] In some examples, in stage 1010, as described above, signal-based fuzz testing is performed on the second signal of stage 1000. In some examples, the first signal is not modified when the second signal is modified. In some examples, as described in conjunction with stage 1000, fuzzer evaluation function block 50 determines which hooks the data unit reaches and / or determines other effects of the second signal.
[0282] In some examples, in stage 1020, a signal-based fuzz test is performed on the first signal of stage 1000 and the second signal of stage 1010. In particular, the fuzz test includes changing bits in positions of the two signals within the data unit. In some examples, as described in conjunction with stage 1000, the fuzzifier evaluation function block 50 determines which hooks the data unit reaches and / or determines other effects of the first signal and the second signal.
[0283] In some examples, in stage 1030, the fuzzer evaluation function block 50 determines whether there is a difference in the effects of: the fuzz test on the first signal of stage 1000 and the fuzz test on the second signal of stage 1010; the combined fuzz test on the first signal and the second signal of stage 1020. For example, if the data unit of stage 1000 reaches a first set of hooks, the data unit of stage 1010 reaches a second set of hooks (which may overlap at least partially with the first set of hooks), and the data unit of stage 1020 reaches a third set of hooks, the fuzzer evaluation function block 50 compares the third set of hooks with the first set of hooks and the second set of hooks. If the third set of hooks contains one or more hooks that are not present in at least one of the first set of hooks (reached by fuzz testing the first signal) and the second set of hooks (reached by fuzz testing the second signal), it is determined that there is a statistical correlation between the two signals in the target binary executable file. If the third set of hooks does not contain any hooks that are not present in at least one of the first set of hooks and the second set of hooks, it is determined that the first signal and the second signal are statistically independent in the target binary executable file.
[0284] In some examples, if the data unit of stage 1020 causes an effect (eg, high CPU load) that is not present in stages 1000 and 1010 , then the first signal and the second signal are determined to be statistically independent in the target binary executable file.
[0285] In some examples, the fuzzifier evaluation function block 50 outputs an indication of the statistical correlation or independence of the first signal of stage 1000 and the second signal of stage 1010. In some examples, the indication is output to a user terminal, such as a user display. In some examples, the indication is stored in a memory. In some examples, a list of signals is stored, and each signal is associated with an indication of its statistical correlation or independence with the other signals.
[0286] In some examples, in stage 1040, the fuzzifier evaluation function block 50 determines whether to perform fuzz testing on the first signal and the second signal together. In some examples, if it is determined in stage 1030 that the first signal and the second signal are statistically independent, the fuzzifier evaluation function block 50 performs signal-based fuzz testing on the first signal and the second signal, respectively. In some examples, in stage 1050, fuzz testing on the first signal and the second signal together is not performed. In some examples, before stage 1040, a signal-based fuzz test is performed on each of the first signal and the second signal, and the determination of stage 1040 is performed only to determine whether to provide further fuzz testing for the combination of the two signals.
[0287] In some examples, separate fuzzy tests are further performed on one or both of the first signal and the second signal in stage 1060. For example, an additional fuzzy test cycle may be performed on the first signal and / or the second signal instead of performing fuzzy testing on the combination of the first signal and the second signal.
[0288] Thus, in some examples, a limited number of data units are used for an initial fuzz testing step of the combined signal.If it is determined that the two signals are not statistically independent, further fuzz testing is performed using additional data units as described above.
[0289] Some examples of the disclosed technology
[0290] Some examples of the above embodiments are listed below. It should be noted that one feature of an isolated example or a combination of one or more features of the example and optionally a combination of one or more features of the example with one or more features of one or more examples below also belong to examples within the scope of the disclosure of the present application.
[0291] Example 1. A system for fuzz testing, the system comprising: a fuzzer data generator configured to continuously generate data units; a first input subsystem configured to input each of the generated data units into a device under test, the input end of the first input subsystem communicating with the output end of the fuzzer data generator, and the output end of the first input subsystem communicating with the device under test; a first fuzz testing agent configured to add each of one or more hooks to a corresponding point of interest in one or more predetermined points of interest in a binary executable file, the binary executable file running on the device under test, wherein in response to the input data unit, each hook outputs information associated with the corresponding point of interest, the output information including data stored in a corresponding address of a memory associated with the corresponding point of interest; and a fuzzer evaluation function block configured to receive information from each of the one or more hooks, wherein the fuzzer data generator communicates with the fuzzer evaluation function block, and the generation of the data unit by the fuzzer data generator is performed in response to the output of the fuzzer evaluation function block.
[0292] Example 2. A system according to any example in this document, in particular Example 1, wherein the fuzz testing agent is embedded in a binary executable file.
[0293] Example 3. A system according to any of the examples in this document, in particular any of Examples 1-2, wherein the first fuzz testing agent is configured to add one or more hooks to the binary executable file without recompiling the binary executable file.
[0294] Example 4. A system according to any of the examples in this document, in particular any of Examples 1-3, wherein, for each of one or more corresponding points of interest, in response to corresponding output information, a fuzzer evaluation function block or a first fuzz testing agent is configured to determine which of the input data units reaches the corresponding hook, and wherein generation of the data unit by the fuzzer data generator is performed in response to a result of the determination.
[0295] Example 5. The system according to any of the examples in this document, in particular Example 4, further includes a timestamp generator, wherein, for each of the input data units, the timestamp generator is configured to set a timestamp associated with inputting the corresponding data unit into the device under test, wherein, for each corresponding point of interest, the timestamp generator is configured to set a corresponding timestamp each time a hook is reached, and wherein determining which of the input data units arrives at the corresponding hook is performed in response to a difference between the timestamp of the corresponding hook and the timestamp of the input data unit.
[0296] Example 6. A system according to any of the examples in this document, in particular Example 4 or 5, wherein, in response to information received at a fuzzer evaluation functional block, the fuzzer evaluation functional block is configured to output an indication of a corresponding point of interest among one or more points of interest to a first fuzz testing agent, and wherein, in response to the output indication of the corresponding point of interest, the first fuzz testing agent is configured to add a corresponding hook to an additional location in the binary executable file associated with the corresponding point of interest.
[0297] Example 7. A system according to any of the examples in this document, in particular Example 6, wherein the fuzzifier evaluation function block is configured to output an indication of the corresponding point of interest to the first fuzz testing agent in response to not receiving information associated with the corresponding point of interest within at least a predetermined time period.
[0298] Example 8. A system according to any example in the present document, in particular example 7 or 8, wherein the additional location is located earlier in the flow of the binary executable file than the corresponding point of interest.
[0299] Example 9. A system according to any of the examples in this document, in particular any of Examples 1-8, wherein, in response to a corresponding hook of one or more hooks not being activated within a predetermined first time period, the fuzzer evaluation function block is configured to: identify a comparison opcode located before the corresponding hook, the comparison opcode having a comparison value and a variable value associated with it; repeatedly receive the comparison value and the variable value from the first fuzz testing agent at multiple instances of the first predetermined time period; control the fuzzer data generator to repeatedly adjust the generated data unit in response to the variable value and the comparison value; and in response to the variable value being equal to the comparison value, determine the necessary adjustment to the generated data unit to make the variable value equal to the comparison value, wherein the fuzzer data generator adjusts the generated data unit according to the necessary adjustment.
[0300] Example 10. A system according to any of the examples in this document, in particular Example 9, wherein the fuzzifier evaluation function block is configured to: repeatedly control or instruct the fuzzifier data generator to insert a predetermined value into a corresponding position of a corresponding data unit, the corresponding position being different at each repetition; and analyze a memory stack associated with a binary executable file to determine which corresponding positions affect the memory stack, and repeatedly adjust the generated data unit in response to the determination result of the corresponding position until the variable value is equal to the comparison value.
[0301] Example 11. A system according to any of the examples in the present document, in particular any of Examples 1-10, wherein the information associated with the corresponding point of interest includes an indication of arrival at the corresponding point of interest, and wherein the fuzzifier evaluation function block is configured to perform a statistical evaluation of the number of times each of one or more predetermined points of interest is initiated.
[0302] Example 12. A system according to any of the examples in this document, in particular any of Examples 1-11, wherein the fuzzifier evaluation function block is configured to compare data stored in a corresponding address of a memory with corresponding data copied from the corresponding address at a previous point in time, and wherein, in response to a result of the comparison indicating that the data is different from the data from the previous point in time, the fuzzifier evaluation function block outputs an indication that there is a difference.
[0303] Example 13. The system according to any of the examples in this document, in particular any of the examples 1-12, further includes: a second fuzz testing agent associated with a copy of the binary executable file running on an emulator or a virtual machine; and a second input subsystem, the second input subsystem configured to input each of the generated data units into the emulator or the virtual machine, the input end of the second input subsystem communicating with the output end of the fuzzer data generator, and the output end of the second input subsystem communicating with the emulator or the virtual machine, wherein a corresponding point of interest among the one or more predetermined points of interest is an entry point of a function, wherein, in response to information received from a hook associated with the entry point of the function, the fuzzer evaluation function block is configured to generate a snapshot of the memory, the snapshot including instructions and values stored in each address from the start of the process of the binary executable file to the entry point of the function, wherein, based at least in part on the generated snapshot, the second fuzz testing agent is configured to set corresponding values of the emulator or the virtual machine so that the data unit input into the emulator or the virtual machine will reach the entry point of the function within the copy of the binary executable file.
[0304] Example 14. A method for fuzz testing, the method comprising: continuously generating data units; inputting each of the generated data units into a device under test; and adding each of one or more hooks to a corresponding point of interest in one or more predetermined points of interest in a binary executable file, the binary executable file running on the device under test, wherein, in response to the input data unit, each hook outputs information associated with the corresponding point of interest, the output information including data stored in a corresponding address of a memory associated with the corresponding point of interest, wherein the generation of the data unit is responsive to the output information associated with the corresponding point of interest.
[0305] Example 15. The method according to any example herein, in particular Example 14, wherein adding one or more hooks to the binary executable file is performed without recompiling the binary executable file.
[0306] Example 16. A method according to any of the examples in this document, in particular Example 14 or 15, wherein, for each of one or more corresponding points of interest, in response to corresponding output information, it is determined which data unit of the input data units reaches the corresponding hook, and wherein the data unit is generated in response to the result of the determination.
[0307] Example 17. The method according to any of the examples in this document, in particular Example 16, further includes: for each of the input data units, setting a timestamp associated with inputting the corresponding data unit into the device under test; and for each corresponding point of interest, setting a corresponding timestamp each time a hook is reached, wherein determining which of the input data units reaches the corresponding hook is performed in response to the difference between the timestamp of the corresponding hook and the timestamp of the input data unit.
[0308] Example 18. The method described in accordance with any of the examples in this document, in particular Example 16 or 17, further includes: in response to output information, outputting an indication of a corresponding point of interest among one or more points of interest; and in response to the output indication of the corresponding point of interest, adding a corresponding hook to an additional location associated with the corresponding point of interest in the binary executable file.
[0309] Example 19. The method according to any of the examples herein, in particular Example 18, further includes: in response to not receiving information associated with the corresponding point of interest within at least a predetermined time period, outputting an indication of the corresponding point of interest.
[0310] Example 20. A method according to any example herein, in particular example 18 or 19, wherein the additional location is located earlier in the flow of the binary executable file than the corresponding point of interest.
[0311] Example 21. The method according to any of the examples in this document, in particular any of the examples 16-20, further includes, in response to a corresponding hook in one or more hooks not being activated within a predetermined first amount in a predetermined time period: identifying a comparison opcode located before the corresponding hook, the comparison opcode having a comparison value and a variable value associated with it; repeatedly receiving the comparison value and the variable value at multiple instances of a predetermined time interval; repeatedly adjusting the generated data unit in response to the variable value and the comparison value; and in response to the variable value being equal to the comparison value, determining the necessary adjustment to the generated data unit so that the variable value is equal to the comparison value, wherein the generated data unit is adjusted according to the necessary adjustment.
[0312] Example 22. The method according to any of the examples in this document, in particular Example 20 or 21, further includes: repeatedly inserting a predetermined value into a corresponding position of a corresponding data unit, wherein the corresponding position is different during each repetition; and analyzing a memory stack associated with a binary executable file to determine which corresponding positions affect the memory stack, and repeatedly adjusting the generated data unit in response to the determination result of the corresponding positions until the variable value is equal to the comparison value.
[0313] Example 23. A method according to any of the examples herein, in particular any of Examples 14-22, wherein the information associated with the corresponding point of interest includes an indication of arrival at the corresponding point of interest, and wherein the method further comprises performing a statistical evaluation of the number of times each of one or more predetermined points of interest is initiated.
[0314] Example 24. The method described in accordance with any of the examples in this document, in particular any of the examples 14-23, further includes: comparing the data stored in the corresponding address of the memory with the corresponding data copied from the corresponding address at a previous point in time; and in response to the result of the comparison indicating that the copied data is different from the copied data from the previous point in time, outputting an indication that there is a difference.
[0315] Example 25. A method according to any of the examples herein, in particular any of Examples 14-24, wherein for each of a plurality of signals, data units are continuously generated to perform signal-based fuzz testing on the device under test.
[0316] Example 26. The method according to any of the examples in this document, in particular Example 25, further includes, for each of multiple signals: determining whether a corresponding point of interest among one or more predetermined points of interest has been reached; and performing further fuzzy testing on the corresponding signal based at least in part on the determination that the corresponding point of interest has been reached.
[0317] Example 27. The method according to any example herein, in particular example 25 or 26, further comprises, for each of the plurality of signals, outputting an indication of one or more points of interest reached by the corresponding data unit.
[0318] It should be understood that certain features of the present invention described in the context of separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features of the present invention described in the context of a single embodiment for the sake of brevity may also be provided individually or in any suitable sub-combination.
[0319] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs. Although methods similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods are described herein.
[0320] All publications, patent applications, patents and other references mentioned herein are incorporated herein by reference in their entirety. In the event of a conflict, the present patent application specification (including definitions) shall prevail. In addition, the materials, methods and examples are illustrative only and are not intended to be limiting.
[0321] Those skilled in the art will recognize that the present invention is not limited to what has been specifically shown and described above. On the contrary, the scope of the present invention is defined by the appended claims, and includes combinations and sub-combinations of the various features described above, as well as changes and modifications to it that would occur to those skilled in the art upon reading the foregoing description.
Claims
1. A system for fuzz testing, the system comprising: a fuzzer data generator configured to continuously generate data units; a first input subsystem configured to input each of the generated data units into a device under test, an input of the first input subsystem being in communication with an output of the fuzzifier data generator, and an output of the first input subsystem being in communication with the device under test; a first fuzz testing agent configured to add each of one or more hooks to a corresponding point of interest in one or more predetermined points of interest in a binary executable file that is executed on the device under test, wherein each hook outputs information associated with the corresponding point of interest in response to an input data unit, the output information comprising data stored in a corresponding address of a memory associated with the corresponding point of interest; and a fuzzer evaluation function configured to receive the information from each of the one or more hooks, The fuzzifier data generator is in communication with the fuzzifier evaluation function block, and the generation of the data unit by the fuzzifier data generator is performed in response to an output of the fuzzifier evaluation function block.
2. The system according to claim 1, wherein: The fuzz testing agent is embedded in the binary executable file.
3. The system according to any one of claims 1 to 2, wherein: The first fuzz testing agent is configured to add the one or more hooks to the binary executable file without recompiling the binary executable file.
4. The system according to any one of claims 1 to 3, wherein: For each of the one or more corresponding points of interest, in response to the corresponding output information, the fuzzer evaluation function block or the first fuzz testing agent is configured to determine which of the input data units reaches the corresponding hook, and The generating of the data unit by the fuzzifier data generator is performed in response to the result of the determination.
5. The system according to claim 4, further comprising a timestamp generator, in, For each of the input data units, the timestamp generator is configured to set a timestamp associated with inputting the corresponding data unit into the device under test, wherein, for each corresponding point of interest, the timestamp generator is configured to set a corresponding timestamp each time a hook is reached, and The determining which data unit of the input data units arrives at the corresponding hook is performed in response to a difference between a timestamp of the corresponding hook and a timestamp of the input data unit.
6. The system according to claim 4 or 5, wherein: In response to the information received at the fuzzer evaluation function, the fuzzer evaluation function is configured to output an indication of a corresponding point of interest of the one or more points of interest to the first fuzz testing agent, and Wherein, in response to the outputted indication of the corresponding point of interest, the first fuzz testing agent is configured to add a corresponding hook to an additional location in the binary executable file associated with the corresponding point of interest.
7. The system according to claim 6, wherein: The fuzzer evaluation function is configured to output the indication of the corresponding point of interest to the first fuzz testing agent in response to not receiving information associated with the corresponding point of interest for at least a predetermined period of time.
8. The system according to claim 7 or 8, wherein: The additional location is located earlier in the flow of the binary executable file than the corresponding point of interest.
9. The system according to any one of claims 1 to 8, wherein: In response to a corresponding hook of the one or more hooks not being activated within a predetermined first time period, the fuzzer evaluation function block is configured to: identifying a comparison opcode preceding a corresponding hook, the comparison opcode having a comparison value and a variable value associated therewith; repeatedly receiving the comparison value and the variable value from the first fuzz testing agent over a plurality of instances of the first predetermined time period; controlling the fuzzifier data generator to repeatedly adjust the generated data unit in response to the variable value and the comparison value; and In response to the variable value being equal to the comparison value, determining a necessary adjustment to the generated data unit to make the variable value equal to the comparison value, Wherein, the fuzzer data generator adjusts the generated data unit according to the necessary adjustment.
10. The system according to claim 9, wherein: The fuzzer evaluation function block is configured as: repeatedly controlling or instructing the fuzzifier data generator to insert a predetermined value into a corresponding position of a corresponding data unit, the corresponding position being different at each repetition; and A memory stack associated with the binary executable file is analyzed to determine which corresponding locations affect the memory stack, and the generated data unit is repeatedly adjusted in response to the results of the determination of the corresponding locations until the variable value is equal to the comparison value.
11. The system according to any one of claims 1 to 10, wherein: The information associated with the corresponding point of interest includes instructions for reaching the corresponding point of interest, and The fuzzer evaluation function block is configured to perform a statistical evaluation on the number of times each of the one or more predetermined points of interest is activated.
12. The system according to any one of claims 1 to 11, wherein: The fuzzer evaluation function block is configured to compare the data stored in the corresponding address of the memory with the corresponding data copied from the corresponding address at a previous point in time, and Wherein, in response to a result of the comparing indicating that the data is different from the data from the previous point in time, the fuzzifier evaluation function block outputs an indication that a difference exists.
13. The system according to any one of claims 1 to 12, further comprising: a second fuzz testing agent associated with a copy of the binary executable file running on the emulator or virtual machine; and a second input subsystem configured to input each of the generated data units into the simulator or virtual machine, an input end of the second input subsystem communicating with an output end of the fuzzer data generator, and an output end of the second input subsystem communicating with the simulator or virtual machine, wherein a corresponding point of interest among the one or more predetermined points of interest is an entry point of a function, wherein, in response to information received from a hook associated with an entry point of the function, the fuzzer evaluation function block is configured to generate a snapshot of the memory, the snapshot comprising instructions and values stored in each address from the start of a process of the binary executable file to the entry point of the function, Wherein, based at least in part on the generated snapshot, the second fuzz testing agent is configured to set corresponding values of the emulator or virtual machine so that a data unit input to the emulator or virtual machine will reach an entry point of the function within the copy of the binary executable file.
14. A method for fuzz testing, the method comprising: Continuously generating data units; inputting each of the generated data units into the device under test; and adding each of the one or more hooks to a corresponding point of interest in one or more predetermined points of interest in a binary executable file that is executed on the device under test, wherein each hook outputs information associated with the corresponding point of interest in response to the input data unit, the output information comprising data stored in a corresponding address of a memory associated with the corresponding point of interest, The data unit is generated in response to the output information associated with the corresponding point of interest.
15. The method according to claim 14, wherein: Adding the one or more hooks to the binary executable file is performed without recompiling the binary executable file.
16. The method according to claim 14 or 15, wherein: For each of the one or more corresponding points of interest, in response to the corresponding output information, determining which data unit of the input data units reaches the corresponding hook, and The generation of the data unit is responsive to a result of the determination.
17. The method according to claim 16, further comprising: for each of the input data units, setting a timestamp associated with inputting the corresponding data unit into the device under test; and For each corresponding point of interest, set the corresponding timestamp each time the hook is reached, The determining which data unit of the input data units arrives at the corresponding hook is performed in response to a difference between a timestamp of the corresponding hook and a timestamp of the input data unit.
18. The method according to claim 16 or 17, further comprising: In response to the output information, outputting an indication of a corresponding point of interest among the one or more points of interest; and Responsive to the outputted indication of the respective point of interest, a respective hook is added to the binary executable file at an additional location associated with the respective point of interest.
19. The method of claim 18, further comprising outputting the indication of the corresponding point of interest in response to not receiving information associated with the corresponding point of interest for at least a predetermined period of time.
20. The method according to claim 18 or 19, wherein: The additional location is located earlier in the flow of the binary executable file than the corresponding point of interest.
21. The method of any one of claims 16-20, further comprising, in response to a corresponding hook of the one or more hooks not being activated within a predetermined first amount of the predetermined time period: identifying a comparison opcode preceding a corresponding hook, the comparison opcode having a comparison value and a variable value associated therewith; repeatedly receiving the comparison value and the variable value at a plurality of instances at predetermined time intervals; repeatedly adjusting the generated data unit in response to the variable value and the comparison value; and In response to the variable value being equal to the comparison value, determining a necessary adjustment to the generated data unit to make the variable value equal to the comparison value, Wherein, the generated data unit is adjusted according to the necessary adjustment.
22. The method according to claim 20 or 21, further comprising: Repeatedly inserting a predetermined value into a corresponding position of a corresponding data unit, wherein the corresponding position is different each time the value is repeated; and A memory stack associated with the binary executable file is analyzed to determine which corresponding locations affect the memory stack, and the generated data unit is repeatedly adjusted in response to the results of the determination of the corresponding locations until the variable value is equal to the comparison value.
23. The method according to any one of claims 14 to 22, wherein: The information associated with the corresponding point of interest includes instructions for reaching the corresponding point of interest, and The method further comprises performing a statistical evaluation of the number of times each of the one or more predetermined points of interest is activated.
24. The method according to any one of claims 14 to 23, further comprising: comparing the data stored in the corresponding address of the memory with corresponding data copied from the corresponding address at a previous point in time; and In response to a result of the comparison indicating that the copied data is different from the copied data from the previous point in time, an indication that a difference exists is output.
25. The method according to any one of claims 14 to 24, wherein: For each signal of a plurality of signals, the data units are continuously generated to perform signal-based fuzz testing on the device under test.
26. The method of claim 25, further comprising, for each of the plurality of signals: determining whether a corresponding point of interest among the one or more predetermined points of interest has been reached; and Based at least in part on the determination that the respective point of interest has been reached, further fuzz testing is performed on the respective signal.
27. The method of claim 25 or 26, further comprising outputting, for each of the plurality of signals, an indication of one or more points of interest reached by the respective data unit.