Integrated fuzzy testing method based on code semantic similarity
Through an integrated fuzz testing method based on code semantic similarity, combined with index amplification and dynamic resource scheduling, the shortcomings of the existing fuzz testing methods in generalization capabilities and resource utilization are solved, and an efficient, migratory and scalable automated vulnerability detection process is realized.
Patent Information
- Application Number
- CN202510617119.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing fuzz testing methods have shortcomings in generalization capabilities and resource utilization, resulting in inefficiency in testing and waste of resources.
The integrated fuzz testing method based on code semantic similarity is adopted. By obtaining the benchmark code base and the code base to be tested, function-level code slicing and encoding processing is performed, and combined with the index amplification method and dynamic resource balance adjustment method, the most suitable fuzzer is intelligently recommended and resources are scheduled dynamically.
The bitmap coverage and detection vulnerabilities of fuzz testing are significantly improved, and more efficient dynamic allocation of computing resources is achieved, testing efficiency is improved and resource overhead is reduced.
Smart Images

Figure CN120144480A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fuzz testing, and particularly relates to an integrated fuzz testing method based on code semantic similarity. Background Art
[0002] With the continuous improvement of the complexity of information systems, automated vulnerability detection technology plays an increasingly crucial role in software security assurance. Fuzzing, as a testing method that automatically generates inputs to trigger abnormal behaviors of programs, is widely used in tasks such as security testing and program verification due to its simplicity and efficiency. Currently, there are various mature fuzz testing tools, such as AFL, QSYM, Angora, RedQueen, etc., which optimize the testing effect in different dimensions, such as improving path coverage, increasing input diversity, or shortening the time to discover defects.
[0003] However, in actual use, there are significant differences in the testing effects of different fuzz testing tools, and their adaptability is limited by various factors such as program structure, input interface complexity, and control flow depth. For example, AFL is good at dealing with high-frequency mutations of shallow logic, while Angora is more suitable for parsing deep-structured code with complex path conditions. For a new program to be tested, it is often difficult to determine which type of fuzzer should be preferred without prior knowledge, nor can the resource allocation strategy of multiple fuzzers be effectively coordinated, resulting in waste of resources of some tools and even a decrease in the overall testing efficiency.
[0004] Therefore, some studies have tried to achieve the joint application of multiple strategies through an ensemble fuzzing framework. However, the current mainstream methods mostly adopt fixed or heuristic fuzzer combination methods, lacking the perception of code features and the quantitative modeling of fuzzer adaptability. At the same time, the resource scheduling mechanism is often statically set and fails to be dynamically optimized according to the real-time performance indicators feedback by the fuzzer during the testing process. These factors limit the portability and adaptability of fuzz testing among different software systems.
[0005] On the other hand, there are a large number of transferable structural features and semantic fragments inside the program, such as common input parsing functions, error handling paths, and state transition logics. These code fragments appear repeatedly in different forms in multiple code libraries and have strong semantic similarity. If a feature representation method based on code semantics can be constructed, the function-level code fragments in different code libraries can be compared, and combined with the testing performance of historical fuzzers in similar structures, it is possible to intelligently recommend the most suitable fuzzer for the new program, thereby improving the testing efficiency and reducing the resource overhead.
[0006] Therefore, there is an urgent need to propose a fuzz testing optimization method based on code semantic representation and historical test experience, which combines function-level code slicing, vectorized modeling, and similarity measurement to establish an intelligent mechanism for fuzzer recommendation and resource scheduling, so as to achieve an efficient, transferable, and scalable automated vulnerability detection process and break through the bottlenecks of existing fuzz testing in terms of generalization ability and resource utilization. Summary of the Invention
[0007] To solve the problems existing in the background technology, the present invention provides a fuzz testing and dynamic resource adjustment method based on code semantic similarity, which solves the technical problems of the deficiencies of existing fuzz testing methods in terms of generalization ability and resource utilization.
[0008] The technical solutions adopted by the present invention include: I. An integrated fuzz testing method based on code semantic similarity S1. Obtain a number of benchmark code libraries, and use multiple fuzz testing tools to comprehensively process each benchmark code library to obtain a number of function-level vectors for each benchmark code library, as well as the total edge coverage rate and index of each fuzz testing tool on the benchmark code library. According to the total edge coverage rate, index, and function-level vectors, obtain the optimal fuzz testing tool for each function-level vector.
[0009] S2. Obtain the code library to be tested, perform function-level code slicing and encoding processing on the code library to be tested to obtain a number of function-level vectors to be tested for the code library to be tested. According to the function-level vectors, the optimal fuzz testing tools for the function-level vectors, and the function-level vectors to be tested, obtain the total scores and indexes of each fuzz testing tool. According to the total edge coverage rate and index, perform index amplification processing on the total scores to obtain the new total scores and indexes of all fuzz testing tools in the code library to be tested.
[0010] S3. Continuously perform fuzz testing on the code library to be tested using the dynamic resource balance adjustment method according to the new total score indexes of all fuzz testing tools, and finally obtain the bitmap coverage rate and the number of detected vulnerabilities of all fuzz testing tools on the code library to be tested.
[0011] The benchmark code libraries are code libraries covering different programming languages, application fields, and code complexities. The code library to be tested includes code libraries in the field of image processing and code libraries in the field of embedded databases, etc. The fuzz testing tools are automated tools for software testing.
[0012] The specific steps of step S1 are as follows: S11. Obtain a number of benchmark code libraries, and use multiple fuzz testing tools to perform fuzz testing on each benchmark code library to obtain the edge coverage rate of each fuzz testing tool for each benchmark code library.
[0013] S12. Use a static analysis tool to perform function-level code slicing on each benchmark code library to obtain several function-level code fragments of each benchmark code library.
[0014] S13. For each fuzzing tool, use the gcov parsing method to obtain the edge coverage rate of each fuzzing tool in each function-level code fragment based on the edge coverage rate of each fuzzing tool for each benchmark code library and the several function-level code fragments of each benchmark code library.
[0015] S14. Statistically calculate the total edge coverage rate of all function-level code fragments of each fuzzing tool, and sort the total edge coverage rates of each fuzzing tool from high to low to obtain the total edge coverage rate index of each fuzzing tool.
[0016] S15. For each function-level code fragment, select the fuzzing tool corresponding to the highest edge coverage rate as the optimal fuzzing tool for the corresponding function-level code fragment.
[0017] S16. Use the UniXcoder model to encode each function-level code fragment to obtain a function-level vector. Furthermore, the optimal fuzzing tool for the function-level code fragment is also used as the optimal fuzzing tool for the corresponding function-level vector.
[0018] The specific steps of step S2 are as follows: S21. Obtain the code library to be tested, and use a static analysis tool to perform function-level code slicing on the code library to be tested to obtain several function-level code fragments to be tested of the code library to be tested.
[0019] S22. Use the UniXcoder model to encode each function-level code fragment to be tested to obtain a function-level vector to be tested.
[0020] S23. Use the cosine similarity method to process each function-level vector to be tested and all function-level vectors of the benchmark code library to obtain the similarity value between each function-level vector to be tested and each function-level vector.
[0021] S24. For each function-level code fragment to be tested, select the highest similarity value from the similarity values between the function-level vector to be tested corresponding to the function-level code fragment to be tested and all function-level vectors as the most similar value of the function-level vector to be tested, and further as the most similar value of the corresponding function-level code fragment to be tested.
[0022] S25. The optimal fuzzing tool for the function-level vector corresponding to the most similar value is used as the optimal fuzzing tool for the function-level vector to be tested. Furthermore, the optimal fuzzing tool for the function-level vector to be tested is also used as the optimal fuzzing tool for the corresponding function-level code fragment to be tested.
[0023] S26. Use the similarity update method based on cyclomatic complexity to process the most similar value of each code snippet at the function level to be tested, and obtain the processed most similar value of each code snippet at the function level to be tested; the processed most similar value of each code snippet at the function level to be tested is used as the score of the corresponding fuzz testing tool.
[0024] The cyclomatic complexity is an index to measure the complexity of a function or code snippet, mainly used to evaluate the structural complexity and testing difficulty of the code.
[0025] S27. For the code library to be tested, count the total scores of each fuzz testing tool in all code snippets at the function level to be tested, that is, obtain the total score of each fuzz testing tool on the code library to be tested, and sort the total scores of each fuzz testing tool from high to low to obtain the total score index of each fuzz testing tool.
[0026] S28. Use the index amplification method to process the total score of the fuzz testing tool on the code library to be tested, and obtain the new total score and index of each fuzz testing tool in the code library to be tested.
[0027] The similarity update method based on cyclomatic complexity in step S26 is specifically as follows: D1. Use the cyclomatic complexity calculation method for each code snippet at the function level to be tested to obtain the cyclomatic complexity of the code snippet at the function level to be tested.
[0028] D2. Process each code snippet at the function level to be tested according to the cyclomatic complexity of all code snippets at the function level to be tested by using the cyclomatic complexity weight adjustment method, and obtain the cyclomatic complexity weight of each code snippet at the function level to be tested.
[0029] D3. For each code snippet at the function level to be tested, take the sum of the product of the most similar value of the code snippet at the function level to be tested and the cyclomatic complexity weight and the most similar value as the processed most similar value of the code snippet at the function level to be tested.
[0030] The cyclomatic complexity weight adjustment method in step D2 is set according to the following formula: W i =ln((f i -min(f)+ε) / (max(f)-min(f)+ε))+f i / max(f) where, W i is the cyclomatic complexity weight of the i-th code snippet at the function level to be tested; f i is the cyclomatic complexity of the i-th code snippet at the function level to be tested; min(f) is the minimum value of the cyclomatic complexity of all code snippets at the function level to be tested; max(f) is the maximum value of the cyclomatic complexity of all code snippets at the function level to be tested; ln( ) represents the natural logarithm; ε is a constant.
[0031] The index amplification method in the step S28 is specifically as follows: Compare the total edge coverage index and the total score index of each fuzz testing tool: If the total score index is ahead of the total edge coverage index, amplify the total score of the fuzz testing tool; if the total score index is behind or unchanged relative to the total edge coverage index, do not process the total score of the fuzz testing tool. After all fuzz testing tools are compared, re - sort the total scores of each fuzz testing tool from high to low according to the amplified total scores to obtain the new total scores and indexes of each fuzz testing tool.
[0032] The amplification process is set according to the following formula: S’ = S(1 + αR’) R’ = R pre -R rec Where, S’ represents the total score of the fuzz testing tool after amplification; S represents the total score of the fuzz testing tool before amplification; α represents the adjustment coefficient; R’ represents the change in the index of the fuzz testing tool; R pre represents the total edge coverage index of the fuzz testing tool; R rec represents the total score index of the fuzz testing tool.
[0033] The step S3 is specifically as follows: S31. Normalize the new total scores of all fuzz testing tools to obtain the new total score percentage of each fuzz testing tool.
[0034] S32. Divide the initial resource allocation ratio of the fuzz testing tools according to the new total score percentage of each fuzz testing tool, and perform fuzz testing on the code library under test according to the initial resource allocation ratio of each fuzz testing tool.
[0035] S33. When the time of fuzz testing reaches the end of the preset time window, use the AFL detection and instrumentation tool for detection to obtain the increase in bitmap coverage and the increase in the number of detected vulnerabilities of each fuzz testing tool respectively.
[0036] S34. Obtain the bitmap coverage and the number of detected vulnerabilities of all fuzz testing tools on the code library under test at the end of the preset time window according to the increase in bitmap coverage and the increase in the number of detected vulnerabilities of each fuzz testing tool.
[0037] S35. Normalize the increase in bitmap coverage at the end of the preset time window of all fuzz testing tools to obtain the increase percentage of bitmap coverage of each fuzz testing tool, and use the increase percentage of bitmap coverage as the contribution rate of the fuzz testing tool.
[0038] S36. At the end of the preset time window, re-adjust the resource allocation ratio of each fuzzing tool according to the dynamic resource balancing adjustment method.
[0039] S37. Repeat steps S33 - S36 until the preset fuzzing termination time is reached, then stop the fuzzing test, and obtain the bitmap coverage rate and the number of detected vulnerabilities of all fuzzing tools on the code library to be tested finally.
[0040] In step S36, each of the fuzzing tools adjusts the resource allocation ratio according to the formula of the following dynamic resource balancing adjustment method: A = Min((B - C) × C, θrh), B ≥ C A = Min((B - C) × C, -θrh), B < C Wherein, A represents the resource allocation ratio that the fuzzing tool needs to adjust; B represents the contribution rate of the fuzzing tool at the end of the preset time window; C represents the resource allocation ratio of the fuzzing tool within the preset time window; θrh represents the preset resource allocation ratio threshold, and Min( ) represents taking the minimum value.
[0041] II. A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0042] III. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0043] The innovation of the present invention lies in proposing an integrated fuzzing method combining an index amplification method and a dynamic resource balancing adjustment method based on code semantic similarity, which improves the bitmap coverage rate of fuzzing and the number of detected vulnerabilities, and realizes a more efficient dynamic allocation of computing resources.
[0044] The beneficial effects of the present invention are: 1. The present invention proposes a new integrated fuzzing method based on code semantic similarity, and realizes a code semantic-driven fuzzer recommendation and resource dynamic scheduling mechanism.
[0045] 2. The index amplification method and the dynamic resource balancing adjustment method proposed by the present invention significantly improve the bitmap coverage rate of fuzzing and realize a more efficient dynamic allocation of computing resources.
[0046] 3. The present invention realizes a fully automated process, reduces the need for manual intervention, and does not require manual intervention from code analysis, recommendation to resource adjustment, which greatly reduces the configuration time compared with traditional methods. Description of the Drawings
[0047] Figure 1 It is a comparison graph of the change curves of the bitmap coverage rates obtained by performing fuzz testing using the method of the present invention and by performing fuzz testing using a single fuzz testing tool.
[0048] Figure 2 It is a comparison graph of the change curves of the number of vulnerabilities obtained by performing fuzz testing using the method of the present invention and by performing fuzz testing using a single fuzz testing tool.
[0049] Figure 3 It is a comparison graph of the change curves of the bitmap coverage rates obtained by the method of the present invention and by the method without performing dynamic resource balance adjustment.
[0050] Figure 4 It is a comparison graph of the change curves of the number of vulnerabilities obtained by the method of the present invention and by the method without performing dynamic resource balance adjustment.
[0051] Figure 5 It is a comparison graph of the change curves of the bitmap coverage rates obtained by the method of the present invention and by the method without performing index magnification.
[0052] Figure 6 It is a comparison graph of the change curves of the number of vulnerabilities obtained by the method of the present invention and by the method without performing index magnification.
[0053] Figure 7 It is a comparison graph of the change curves of the bitmap coverage rates obtained by the method of the present invention and by the existing integrated fuzz testing method Autofz.
[0054] Figure 8 It is a comparison graph of the change curves of the number of vulnerabilities obtained by the method of the present invention and by the existing integrated fuzz testing method Autofz. Specific embodiments
[0055] The following further elaborates on the present invention in conjunction with the accompanying drawings and embodiments. However, the present invention is not limited thereto. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also considered within the protection scope of the present invention. The content not described in detail in this specification belongs to the prior art well-known to those of ordinary skill in the art.
[0056] The specific embodiments of the present invention are as follows: Embodiment 1
[0057] The integrated fuzz testing method of this embodiment includes the following steps: S1. Obtain a number of benchmark code libraries through the data link of the computer and store the benchmark code libraries in the data hard disk. Use the central processing unit to perform fuzz testing on each benchmark code library in turn with multiple fuzz testing tools, and perform function-level code slicing to obtain a number of function-level vectors for each benchmark code library. Obtain the total edge coverage rate and index of each fuzz testing tool according to all function-level vectors, and obtain the optimal fuzz testing tool for each function-level vector according to the total edge coverage rate and index.
[0058] S11. Obtain a number of benchmark code libraries through the data link of the computer and store the benchmark code libraries in the data hard disk. Use the central processing unit to perform fuzz testing on each benchmark code library with multiple fuzz testing tools to obtain the edge coverage rate of each fuzz testing tool for each benchmark code library.
[0059] In specific implementation, the fuzz testing tools include afl, aflfast, angora, fairfuzz, laf-intel, learn afl, mopt, qsym, radamsa, and redqeen.
[0060] The edge coverage rate is a fine-grained code coverage measurement method in fuzz testing.
[0061] S12. Use a static analysis tool to perform function-level code slicing on each benchmark code library to obtain a number of function-level code segments for each benchmark code library.
[0062] In specific implementation, the static analysis tool uses the LLVM parser. The function-level code segment is a code block extracted from the program in units of functions.
[0063] S13. According to the edge coverage rate of each fuzz testing tool for each benchmark code library and the number of function-level code segments of each benchmark code library, use the gcov parsing method to obtain the edge coverage rate of each fuzz testing tool in each function-level code segment.
[0064] S14. Statistically calculate the total edge coverage rate of all function-level code segments of each fuzz testing tool, and sort the total edge coverage rates of each fuzz testing tool from high to low to obtain the total edge coverage rate index of each fuzz testing tool.
[0065] S15. For each function-level code segment, select the fuzz testing tool corresponding to the highest edge coverage rate as the optimal fuzz testing tool for the corresponding function-level code segment.
[0066] Further, a preset edge coverage rate threshold is set. If the highest edge coverage rate is less than the preset edge coverage rate threshold, the function-level code snippet and the corresponding fuzz testing tool are discarded; if it is not less than, the function-level code snippet and the corresponding fuzz testing tool are retained. All discarded function-level code snippets do not execute step S16.
[0067] S16. Use the UniXcoder model to encode each function-level code snippet to obtain a function-level vector. Furthermore, the optimal fuzz testing tool for the function-level code snippet is also used as the optimal fuzz testing tool for the corresponding function-level vector.
[0068] The vector dimension is 768 dimensions × n dimensions, and n is related to the length of the function code snippet.
[0069] S2. Obtain the code library to be tested through the data link of the computer and store the code library to be tested in the data hard disk. Use the central processing unit to perform function-level code slicing and encoding processing on the code library to be tested to obtain several function-level vectors to be tested for the code library to be tested. Obtain the total score and index of each fuzz testing tool according to the function-level vector, the optimal fuzz testing tool of the function-level vector, and the function-level vector to be tested. Perform index amplification processing on the total score according to the total edge coverage rate and index of the fuzz testing tool to obtain the new total score and index of all fuzz testing tools in the code library to be tested.
[0070] S21. Obtain the code library to be tested through the data link of the computer and store the code library to be tested in the data hard disk. Use the central processing unit to perform function-level code slicing on the code library to be tested using a static analysis tool to obtain several function-level code snippets to be tested for the code library to be tested.
[0071] S22. Use the UniXcoder model to encode each function-level code snippet to be tested to obtain a function-level vector to be tested.
[0072] S23. Process each function-level vector to be tested and all function-level vectors of the reference code library using the cosine similarity method to obtain the similarity values between each function-level vector to be tested and each function-level vector of the reference code library.
[0073] The cosine similarity method is set according to the following formula: sim(u,v)=(u·v) / (||u||·||v||) Among them, sim(u, v) represents the similarity between u and v; u represents the function-level vector of the code library to be tested; v represents the function-level vector of the benchmark code library; ||u|| represents the norm of the function-level vector u to be tested; ||v|| represents the norm of the function-level vector v. In specific implementation, the value of sim(u, v) is in the interval [-1, 1]. The closer the value is to 1, the higher the semantic similarity; the closer the value is to -1, the lower the semantic similarity.
[0074] S24. For each function-level code snippet to be tested, select the highest similarity value from the similarity values between the function-level vector corresponding to the function-level code snippet to be tested and all function-level vectors of the benchmark code library as the most similar value of the function-level vector to be tested, and further as the most similar value of the corresponding function-level code snippet to be tested.
[0075] S25. The optimal fuzz testing tool of the function-level vector of the benchmark code library corresponding to the most similar value is used as the optimal fuzz testing tool of the function-level vector to be tested corresponding to the most similar value, and further the optimal fuzz testing tool of the function-level vector to be tested is also used as the optimal fuzz testing tool of the corresponding function-level code snippet to be tested.
[0076] S26. Process the most similar value of each function-level code snippet to be tested by using a similarity update method based on cyclomatic complexity to obtain the processed most similar value of each function-level code snippet; the processed most similar value of each function-level code snippet is used as the score of the corresponding fuzz testing tool.
[0077] The cyclomatic complexity is an index to measure the complexity of a function or code snippet, mainly used to evaluate the structural complexity and testing difficulty of the code.
[0078] The similarity update method based on cyclomatic complexity is specifically as follows: D1. Use the cyclomatic complexity calculation method for each function-level code snippet to be tested to obtain the cyclomatic complexity of the function-level code snippet to be tested.
[0079] In specific implementation, the cyclomatic complexity calculation method is set according to the following formula: CCN = E - N + 2P Among them, CCN represents the cyclomatic complexity of the function-level code snippet; E represents the total number of paths of process transfer in the function-level code snippet, and also represents the number of edges in the control flow graph after the function-level code snippet is transformed into a control flow graph; N represents the total number of points such as program blocks or conditional judgments in the function-level code snippet, and also represents the number of nodes in the control flow graph after the function-level code snippet is transformed into a control flow graph; P represents the number of function-level code snippets. Here, it is for the number of function-level code snippets themselves, so it is 1.
[0080] D2. Process each code snippet at the function level to be tested using the cyclomatic complexity weight adjustment method based on the cyclomatic complexity of all code snippets at the function level to be tested, and obtain the cyclomatic complexity weight of each code snippet at the function level to be tested.
[0081] The cyclomatic complexity weight adjustment method in step D2 is set according to the following formula: W i = ln((f i - min(f) + ε) / (max(f) - min(f) + ε)) + f i / max(f) where W i is the cyclomatic complexity weight of the i-th code snippet at the function level to be tested; f i is the cyclomatic complexity of the i-th code snippet at the function level to be tested; min(f) is the minimum value of the cyclomatic complexity of all code snippets at the function level to be tested; max(f) is the maximum value of the cyclomatic complexity of all code snippets at the function level to be tested; ln( ) represents the natural logarithm, the logarithmic function with the natural constant e as the base; ε is a small constant, which is set to 10 -6 in specific implementations to prevent division by zero.
[0082] D3. For each code snippet at the function level to be tested, use the product of the most similar value of the code snippet at the function level to be tested and the cyclomatic complexity weight plus the most similar value as the most similar value of the code snippet at the function level to be tested after processing.
[0083] The most similar value of each code snippet at the function level to be tested is processed according to the following formula to obtain the updated most similar value after processing: Sim (i)max ’ = (1 + W i ) Sim (i)max where Sim (i)max ’ represents the most similar value of the i-th code snippet at the function level to be tested after processing; Sim (i)max represents the most similar value of the i-th code snippet at the function level to be tested; W i represents the cyclomatic complexity weight of the i-th code snippet at the function level to be tested.
[0084] S27. For the code library to be tested, count the total score of each fuzz testing tool in all code snippets at the function level to be tested, that is, obtain the total score of each fuzz testing tool on the code library to be tested, and sort the total scores of each fuzz testing tool from high to low to obtain the total score index of each fuzz testing tool.
[0085] S28. To highlight the advantages of fuzz testing tools that may perform well in the code library under test but have a low index in the benchmark test, the index amplification method is used to process the total scores of fuzz testing tools on the code library under test, obtaining the new total scores and indexes of each fuzz testing tool in the code library under test.
[0086] The index amplification method is specifically as follows: Compare the total edge coverage index and the total score index of each fuzz testing tool. If the total score index is ahead of the total edge coverage index, perform amplification processing on the total score of the fuzz testing tool. If the total score index is behind or unchanged relative to the total edge coverage index, do not process the total score of the fuzz testing tool. After comparing all fuzz testing tools, re - sort them from high to low according to the total scores of each fuzz testing tool after amplification processing to obtain the new total scores and indexes of each fuzz testing tool.
[0087] The amplification processing is set according to the following formula: S’ = S(1 + αR’) R’ = R pre -R rec Among them, S’ represents the total score of the fuzz testing tool after amplification processing; S represents the total score of the fuzz testing tool before amplification processing; α represents the adjustment coefficient; R’ represents the change in the index of the fuzz testing tool; R pre represents the total edge coverage index of the fuzz testing tool; R rec represents the total score index of the fuzz testing tool.
[0088] S3. Continuously perform fuzz testing on the code library under test using the dynamic resource balance adjustment method according to the new total score indexes of all fuzz testing tools, and finally obtain the bitmap coverage rate and the number of detected vulnerabilities of all fuzz testing tools on the code library under test.
[0089] The benchmark code library is a code library covering different programming languages, application fields, and code complexities. The code library under test includes code libraries in the field of image processing and code libraries in the field of embedded databases, etc.
[0090] The bitmap coverage rate is a measurement method for measuring the code path coverage during the fuzz testing process.
[0091] S31. Normalize the new total scores of all fuzz testing tools to obtain the new total score percentage of each fuzz testing tool.
[0092] S32. Divide the initial resource allocation ratio of fuzz testing tools according to the new total score percentage of each fuzz testing tool, and perform fuzz testing on the code library under test according to the initial resource allocation ratio of each fuzz testing tool.
[0093] S33. When the time of fuzz testing reaches the end of the preset time window, use the AFL detection and instrumentation tool to detect, and obtain the increase in bitmap coverage rate and the increase in the number of detected vulnerabilities of each fuzz testing tool respectively.
[0094] S34. Based on the increase in bitmap coverage rate and the number of detected vulnerabilities of each fuzz testing tool, obtain the bitmap coverage rate and the number of detected vulnerabilities of all fuzz testing tools on the code library to be tested at the end of the preset time window.
[0095] S35. Normalize the increase in bitmap coverage rate of all fuzz testing tools at the end of the preset time window to obtain the percentage increase in bitmap coverage rate of each fuzz testing tool, and use the percentage increase in bitmap coverage rate as the contribution rate of the fuzz testing tool.
[0096] S36. When reaching the end of the preset time window, readjust the resource allocation ratio of each fuzz testing tool according to the dynamic resource balance adjustment method.
[0097] In step S36, the resource allocation ratio of each fuzz testing tool is adjusted according to the following formula of the dynamic resource balance adjustment method: A = Min((B - C) × C, θrh), B ≥ C A = Min((B - C) × C, -θrh), B < C Wherein, A represents the resource allocation ratio that needs to be adjusted for the fuzz testing tool; B represents the contribution rate of the fuzz testing tool at the end of the preset time window; C represents the resource allocation ratio of the fuzz testing tool within the preset time window; θrh represents the preset resource allocation ratio threshold, and Min( ) represents taking the minimum value.
[0098] In specific implementation, if the new total score percentage of one of the fuzz testing tools is 30%, the initial resource allocation ratio is 30% CPU resources, the preset resource allocation ratio threshold θrh is 5%, the preset time window is 5 minutes, and the percentage increase in bitmap coverage rate of the fuzz testing tool within the first preset time window is 25%, that is, the contribution rate of the fuzz testing tool within the first preset time window is 25%. Then, the resource allocation ratio that needs to be adjusted for this fuzz testing tool is: A = Min((25% - 30%) × 30%, -5%) = -1.5%. Therefore, at the end of the preset time window, the resource allocation ratio of this fuzz testing tool is decreased by 1.5%. If the result of A is negative, it means that the resource allocation ratio of this fuzz testing tool is decreased; if the result of A is positive, it means that the resource allocation ratio of this fuzz testing tool is increased.
[0099] S37. Repeat steps S33 - S36 until the preset fuzz testing termination time is reached, then stop the fuzz testing and obtain the bitmap coverage rate and the number of detected vulnerabilities of all fuzz testing tools on the code library under test finally.
[0100] The innovation of the present invention lies in proposing an integrated fuzz testing method that combines an index amplification method and a dynamic resource balance adjustment method based on code semantic similarity, improving the bitmap coverage rate of fuzz testing and the number of detected vulnerabilities, and achieving a more efficient dynamic allocation of computing resources.
[0101] Embodiment 2
[0102] The benchmark code libraries used in this embodiment include: boringssl, muparser, bzip, openssl101f, eigen, openssl102d, freetype2, openssl110c, guetzli, openh264, haproxy, openthread, harfbuzz, pcre2, jbig2dec, phmap, libarchive, proj4, libevent, quickjs, libexif, re2, libgd, spdlog, libjpeg, sqlite, libpng, tidy - html, libraw, unit, libssh, uriparser, libunwind, vorbis, libxml2, woff2, lcms, zlib, libyaml, and zopfli, a total of 40 benchmark code libraries. The code library under test uses: exiv2.
[0103] Among them, boringssl, openssl101f, openssl102d, openssl110c, libssh, and libyaml are code libraries in the field of encryption and security; libarchive, libexif, libxml2, and tidy-html are code libraries in the field of file parsing; libjpeg, libpng, libraw, guetzli, lcms, jbig2dec, freetype2, harfbuzz, libgd, exiv2, and woff2 are code libraries in the field of image processing; haproxy, libevent, and openthread are code libraries in the field of network protocols; sqlite is a code library in the field of databases; vorbis andopenh264 are code libraries in the field of multimedia processing; quickjs and unit are code libraries in the field of web development; pcre2, re2, spdlog, and uriparser are code libraries in the field of text processing; bzip2, zlib, and zopfli are code libraries in the field of compression algorithms; eigen and proj4 are code libraries in the field of mathematics and scientific computing; muparser is a code library for parsing mathematical expressions; libunwind is a code library for handling function call stacks; and phmap is a high-performance hash table code library.
[0104] The fuzz testing tools used in this embodiment include: afl, aflfast, angora, fairfuzz, laf-intel, learn-afl, mopt, qsym, radamsa, and redqeen, a total of 10 fuzz testing tools.
[0105] This embodiment is implemented using the same method steps as in Embodiment 1.
[0106] During the implementation process, a static analysis tool is used to perform function-level code slicing on each benchmark code library, and a total of 36,734 function-level code snippets are obtained.
[0107] During the implementation process, the various fuzz testing tools obtained through statistics are sorted from high to low according to the total edge coverage rate as shown in Table 1.
[0108] Table 1 Fuzzing Tool Total Edge Coverage Total Edge Coverage Index redqueen 82.1% 1 lafintel 69.6% 2 qsym 64.7% 3 fairfuzz 59.5% 4 learnafl 58.9% 5 mopt 58.8% 6 angora 57.8% 7 radamsa 55.9% 8 afl 55.3% 9 aflfast 55.0% 10 During the implementation process, the various fuzz testing tools obtained through statistics are sorted from high to low according to the total score as shown in Table 2.
[0109] Table 2 Fuzzing Tool Total Score Total Score Index redqueen 1218.205337 1 lafintel 1006.198804 2 qsym 911.487161 3 learnafl 898.780717 4 angora 861.809616 5 mopt 858.846206 6 fairfuzz 834.319373 7 radamsa 819.102313 8 afl 800.735747 9 aflfast 764.344617 10 During the implementation process, after using the index amplification method, the statistically obtained fuzz testing tools are sorted from high to low according to the new total scores as shown in Table 3.
[0110] Table 3 Fuzzing Tool New Total Score New Total Score Index redqueen 1218.205337 1 lafintel 1006.198804 2 angora 930.754385 3 qsym 911.487161 4 learnafl 898.780717 5 fairfuzz 858.846206 6 mopt 834.319373 7 radamsa 819.102313 8 afl 800.735747 9 aflfast 764.344617 10 According to the new total scores of each fuzz testing tool in Table 3, the new total scores are normalized, and the new total score percentages of each fuzz testing tool are redqueen: 14.18%, lafintel: 10.34%, angora: 12.53%, qsym: 9.68%, learnafl: 9.38%, radamsa: 8.93%, mopt: 8.86%, fairfuzz: 9.01%, afl: 8.64%, and aflfast: 8.45%.
[0111] Initially, the resource allocation ratio is divided according to the new total score percentage obtained above, and the fuzzy test is performed on the code base to be tested according to the obtained initial resource allocation ratio.
[0112] The dynamic resource balancing adjustment method proposed in the present invention is used to readjust the resource allocation ratio of each fuzz testing tool until the preset fuzz testing termination time is reached, then the fuzz testing is stopped. Finally, the total bitmap coverage and total number of vulnerabilities that can be achieved by fuzz testing the code library to be tested are obtained based on the bitmap coverage and number of vulnerabilities of all fuzz testing tools.
[0113] The variation curves of the bitmap coverage and time of all fuzz testing tools implemented in this embodiment on the code base to be tested are as follows: Figure 1 shown.
[0114] Comparative Example 1 In this comparative example, redqueen, lafintel, angora, qsym, learnafl, radamsa, mopt, fairfuzz, afl and aflfast fuzz testing tools are used to perform fuzz testing on the exiv2 code base to be tested. The change curves of the bitmap coverage and time of each fuzz testing tool on the code base to be tested are shown in the figure. Figure 1 The curves of the number of vulnerabilities detected by each fuzz testing tool on the code base under test and the change of time are shown in Figure 2 As shown. Figure 1 and Figure 2 It can be seen that the method of the present invention has great advantages over the conventional use of a single fuzz testing tool to perform fuzz testing on the code base to be tested.
[0115] Comparative Example 2 This comparative example was implemented using the same steps as in Example 2. However, in this comparative example, after obtaining the initial resource allocation ratios of each fuzz testing tool, fuzz testing was directly carried out according to the initial resource allocation ratios, and no dynamic resource balance adjustment was performed. Finally, the curves of the bitmap coverage rate and time of all fuzz testing tools on the code library to be tested are as follows Figure 3 shown, and the curves of the number of vulnerabilities detected by all fuzz testing tools on the code library to be tested and time are as follows Figure 4 shown. From Figure 3 and Figure 4 , it can be seen that the dynamic resource balance adjustment method adopted by the method of the present invention significantly improves the performance during fuzz testing.
[0116] Comparative Example 3 This comparative example was implemented using the same steps as in Example 2. However, in this comparative example, after obtaining the total score indexes of each fuzz testing tool, the index amplification method was not used for processing. Instead, normalization, division of resource allocation ratios, and dynamic resource balance adjustment processing were directly carried out according to the total scores without index amplification. Finally, the curves of the bitmap coverage rate and time of all fuzz testing tools on the code library to be tested are as follows Figure 5 shown, and the curves of the number of vulnerabilities detected by all fuzz testing tools on the code library to be tested and time are as follows Figure 6 shown. From Figure 5 and Figure 6 , it can be seen that the index amplification method adopted by the method of the present invention also significantly improves the performance during fuzz testing.
[0117] Comparative Example 4 This comparative example used the existing integrated fuzz testing method Autofz to perform fuzz testing on the exiv2 code library to be tested. The curves of the bitmap coverage rate and time obtained from the fuzz testing are as follows Figure 7 shown, and the curves of the number of vulnerabilities detected by the fuzz testing and time are as follows Figure 8 shown. From Figure 7 and Figure 8 , it can be seen that the method of the present invention has better effects than the existing integrated fuzz testing method Autofz in terms of the bitmap coverage rate and the number of detected vulnerabilities.
[0118] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can also make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. An integrated fuzz testing method based on code semantic similarity, characterized in that: The method comprises the following steps: S1. Obtain several benchmark code bases, use multiple fuzz testing tools to perform comprehensive processing on each benchmark code base to obtain several function-level vectors of each benchmark code base and the total edge coverage and index of each fuzz testing tool on the benchmark code base, and obtain the optimal fuzz testing tool for each function-level vector according to the total edge coverage and index and function-level vector; S2. Obtain a code base to be tested, perform function-level code slicing and encoding processing on the code base to be tested to obtain several function-level vectors to be tested of the code base to be tested, obtain the total score and index of each fuzz testing tool according to the function-level vector, the optimal fuzz testing tool of the function-level vector and the function-level vector to be tested, and perform index amplification processing on the total score according to the total edge coverage and index to obtain a new total score and index of all fuzz testing tools in the code base to be tested; S3. Based on the new total score index of all fuzz testing tools, a dynamic resource balancing adjustment method is used to continuously perform fuzz testing on the code base to be tested, and finally the bitmap coverage and the number of detected vulnerabilities of all fuzz testing tools on the code base to be tested are obtained.
2. The integrated fuzz testing method according to claim 1, characterized in that: The step S1 is specifically as follows: S11, obtaining several benchmark code bases, using multiple fuzz testing tools to perform fuzz testing on each benchmark code base, and obtaining the edge coverage of each fuzz testing tool on each benchmark code base; S12, using a static analysis tool to perform function-level code slicing on each benchmark code library to obtain a number of function-level code snippets for each benchmark code library; S13, according to the edge coverage of each fuzz testing tool on each benchmark code base and several function-level code snippets of each benchmark code base, using the gcov parsing method to obtain the edge coverage of each fuzz testing tool on each function-level code snippet; S14, counting the total edge coverage of all function-level code snippets of each fuzz testing tool, and sorting the total edge coverage of each fuzz testing tool from high to low to obtain a total edge coverage index of each fuzz testing tool; S15. For each function-level code snippet, select a fuzz testing tool corresponding to the highest edge coverage as the optimal fuzz testing tool for the corresponding function-level code snippet; S16. Use the UniXcoder model to encode each function-level code snippet to obtain a function-level vector, and then the optimal fuzz testing tool for the function-level code snippet is also used as the optimal fuzz testing tool for the corresponding function-level vector.
3. The integrated fuzz testing method according to claim 1, characterized in that: The step S2 is specifically as follows: S21. Obtain a code base to be tested, and use a static analysis tool to perform function-level code slicing on the code base to be tested to obtain several function-level code fragments to be tested of the code base to be tested; S22, using the UniXcoder model to encode each function-level code snippet to be tested to obtain a function-level vector to be tested; S23, processing each function-level vector to be tested and all function-level vectors of the benchmark code library using a cosine similarity method to obtain a similarity value between each function-level vector to be tested and each function-level vector; S24, for each function-level code snippet to be tested, selecting the highest similarity value from the similarity values of the function-level vector to be tested corresponding to the function-level code snippet to be tested and all function-level vectors as the most similarity value of the function-level vector to be tested, and then as the most similarity value of the corresponding function-level code snippet to be tested; S25, the optimal fuzz testing tool of the function-level vector corresponding to the most similarity value is used as the optimal fuzz testing tool of the function-level vector to be tested, and then the optimal fuzz testing tool of the function-level vector to be tested is also used as the optimal fuzz testing tool of the corresponding function-level code fragment to be tested; S26, using a similarity update method based on cyclomatic complexity to process the most similarity value of each function-level code snippet to be tested, to obtain the most similarity value of each function-level code snippet to be tested after processing; the most similarity value of each function-level code snippet to be tested after processing is used as the score of the corresponding fuzz testing tool; S27. For the code base to be tested, the total score of each fuzz testing tool in all function-level code snippets to be tested is counted, that is, the total score of each fuzz testing tool on the code base to be tested is obtained, and the total scores of each fuzz testing tool are sorted from high to low to obtain the total score index of each fuzz testing tool; S28. Use the index amplification method to process the total score of the fuzz testing tool on the code base to be tested, and obtain a new total score and index of each fuzz testing tool in the code base to be tested.
4. The integrated fuzz testing method according to claim 3, characterized in that: The similarity updating method based on cyclomatic complexity in step S26 is specifically as follows: D1. Using the cyclomatic complexity calculation method for each function-level code snippet to be tested, the cyclomatic complexity of the function-level code snippet to be tested is obtained; D2. According to the cyclomatic complexity of all the function-level code snippets to be tested, a cyclomatic complexity weight adjustment method is used to process each function-level code snippet to be tested, so as to obtain a cyclomatic complexity weight of each function-level code snippet to be tested; D3. For each function-level code snippet to be tested, the product of the most similarity value of the function-level code snippet to be tested and the cyclomatic complexity weight plus the most similarity value is taken as the most similarity value of the function-level code snippet to be tested after processing.
5. The integrated fuzz testing method according to claim 4 is characterized in that: The cyclomatic complexity weight adjustment method in step D2 is set according to the following formula: W i =ln((f i -min(f)+ε) / (max(f)-min(f)+ε))+f i / max(f) Among them, W i is the cyclomatic complexity weight of the i-th function-level code snippet to be tested; f i is the cyclomatic complexity of the i-th function-level code snippet to be tested; min(f) is the minimum cyclomatic complexity of all function-level code snippets to be tested; max(f) is the maximum cyclomatic complexity of all function-level code snippets to be tested; ln( ) represents the natural logarithm; ε is a constant.
6. The integrated fuzz testing method according to claim 3, characterized in that: The index enlargement method in step S28 is specifically as follows: Compare the total edge coverage index and total score index of each fuzz testing tool: if the total score index is higher than the total edge coverage index, the total score of the fuzz testing tool is amplified; if the total score index is lower than the total edge coverage index or remains unchanged, the total score of the fuzz testing tool is not processed; After all fuzz testing tools are compared, they are re-sorted from high to low according to their total scores after amplification to obtain new total scores and indexes of each fuzz testing tool; The amplification process is set according to the following formula: S'=S(1+αR') R’=R pre -R rec Among them, S' represents the total score after the fuzz test tool is amplified; S represents the total score before the fuzz test tool is amplified; α represents the adjustment coefficient; R' represents the change of the fuzz test tool index; R pre represents the total edge coverage index of the fuzz testing tool; R rec Represents the total score index of the fuzz testing tool.
7. The integrated fuzz testing method according to claim 3, characterized in that: The step S3 is specifically as follows: S31, normalizing the new total scores of all fuzz testing tools to obtain the new total score percentage of each fuzz testing tool; S32, dividing the initial resource allocation ratio of the fuzz testing tools according to the new total score percentage of each fuzz testing tool, and performing fuzz testing on the code base to be tested according to the initial resource allocation ratio of each fuzz testing tool; S33, when the fuzz test time reaches the end of a preset time window, the AFL detection instrumentation tool is used to perform detection, and the increase in the bitmap coverage of each fuzz test tool and the increase in the number of detected vulnerabilities are obtained respectively; S34, obtaining the bitmap coverage and the amount of detected vulnerabilities of all fuzz testing tools on the code base to be tested at the end of the preset time window according to the bitmap coverage and the increase in the amount of vulnerabilities of each fuzz testing tool; S35, normalizing the increase in bitmap coverage at the end of the preset time window of all fuzz testing tools to obtain the increase percentage of the bitmap coverage of each fuzz testing tool, and taking the increase percentage of the bitmap coverage as the contribution rate of the fuzz testing tool; S36, when the preset time window ends, the resource allocation ratio of each fuzz testing tool is readjusted and allocated according to the dynamic resource balancing adjustment method; S37. Repeat steps S33 to S36 until the preset fuzz test termination time is reached, then stop the fuzz test and obtain the final bitmap coverage and the number of detected vulnerabilities of all fuzz test tools on the code base to be tested.
8. The integrated fuzz testing method according to claim 7, characterized in that: In step S36, each of the fuzzy testing tools adjusts the resource allocation ratio according to the formula of the following dynamic resource balancing adjustment method: A=Min((BC)×C,θrh),B≥C A=Min((BC)×C,-θrh),B<C Among them, A represents the resource allocation ratio that needs to be adjusted for the fuzz testing tool; B represents the contribution rate of the fuzz testing tool at the end of the preset time window; C represents the resource allocation ratio of the fuzz testing tool within the preset time window; θrh represents the preset adjusted resource allocation ratio threshold, and Min ( ) represents the minimum.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Integrated fuzz testing method and system for automatically selecting fuzzer combination
CN116010281A
Fuzzy test method and device, electronic equipment and readable storage medium
CN117828609A
Intelligent contract vulnerability detection method and system based on constraint guide fuzzy test
CN118760606A
Fractional-order MEMS gyroscope acceleration adaptive backstepping control method without accurate reference trajectory
GB202019112D0
Creating an optimal test suite
US20250036554A1