Function fine granularity similarity detection method and system based on sensitive slicing
Through a fine-grained function similarity detection method based on sensitive slicing, utilizing the twin network and Structure2vec model, the serious problems of missed detection and false positives in existing technologies are solved, and higher-precision vulnerability detection is achieved.
Patent Information
- Application Number
- CN202510778122.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-23
Smart Images

Figure CN120688057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security, and in particular to a method and system for detecting fine-grained similarity of functions based on sensitive slicing. Background Art
[0002] With the development of the Internet of Things, embedded devices have become increasingly used in all aspects of social production and life. At the same time, various attacks targeting embedded devices are constantly emerging, and most of these attacks are inseparable from the exploitation of device firmware vulnerabilities. Timely and accurate detection of security vulnerabilities in embedded device firmware has become a key method for mitigating embedded device security threats, and many researchers have devoted extensive work to this area. Because embedded device firmware currently generally adopts code reuse to minimize device development cycles and costs, similarity detection methods based on vulnerability signatures have become a very important technical approach in the field of embedded device firmware vulnerability analysis and mining, and have been proven effective in many previous studies.
[0003] Vulnerability analysis and discovery methods based on known vulnerability features focus on two key aspects: first, characterizing known vulnerabilities; and second, determining whether the features of two objects are consistent. Currently, the latest research on the latter aspect is primarily based on artificial intelligence technologies such as deep neural networks and natural language processing (NLP). However, regardless of the AI technology employed, the former (characterization) serves as a crucial foundation and prerequisite for the latter. The accuracy of feature characterization has become a key factor influencing the ultimate effectiveness of these methods. Existing vulnerability signature models typically analyze the function where the vulnerability resides.
[0004] When constructing features for vulnerable functions, researchers generally choose the entire function as the analysis object, extracting relevant instruction-level features and code semantic features as the overall features of the function. This type of feature model can be used to detect two functions with similar features. In vulnerability detection tasks, if a function is known to contain a vulnerability, it can also be predicted that similar functions may also have vulnerabilities. Whether the vulnerability actually exists requires further analysis and judgment. The essence of this type of feature characterization method is to transform the analysis of target objects such as vulnerabilities into an analysis of the functions in which they are located, and to represent the characteristics of the vulnerability with the characteristics of the vulnerability function. Although finding similar functions through the characteristics of the vulnerability function can achieve a greater probability of discovering vulnerabilities, this type of method currently still faces serious problems of missed reports and false positives. Summary of the Invention
[0005] In view of the fact that existing vulnerability detection methods use the entire vulnerability function as a feature for detection, which still leads to serious problems of missed reports and false alarms, the present invention proposes a fine-grained function similarity detection method and system based on sensitive slices. By extracting the sensitive slices of the function, the features of the sensitive slices are used as the features of the corresponding function for similarity detection, effectively improving the detection accuracy and precision of function similarity detection when applied to scenarios such as function vulnerability discovery.
[0006] In a first aspect, the present invention provides a method for fine-grained function similarity detection based on sensitive slices, comprising:
[0007] Step 1: Build a function sample database and obtain the sensitive slices of each function in the function sample database, and train a function similarity detection model based on the sensitive slices; wherein the function similarity detection model is a twin network structure including two identical sensitive slice embedding generation modules, and the sensitive slice embedding generation module includes Structure2vec and a fully connected layer;
[0008] Step 2: Perform similarity detection on the target function based on the function similarity detection model.
[0009] Furthermore, the step 1 includes:
[0010] Step 1.1: Build a function sample database based on open source software or firmware;
[0011] Step 1.2: Identify sensitive instructions of the functions in the function sample database, and extract sensitive slices corresponding to the functions in the function sample database based on the sensitive instructions; the sensitive instructions include library function call instructions, specific operations or operation sequences; the specific operations or operation sequences are specific operations or operation sequences with known vulnerabilities;
[0012] Step 1.3: determining original features corresponding to sensitive slices of the function in the function sample database; wherein the original features include basic block statistical features, call features, and structural features;
[0013] Step 1.4: training the function similarity detection model based on the original features of the functions in the function sample database.
[0014] Furthermore, the step 1.1 includes:
[0015] Step 1.1.1: Collect open source software or firmware datasets, generate binary programs for different architectures through cross-compilation, and set different optimization levels during compilation. For each architecture, multiple binary programs with different compilation optimization levels can be generated.
[0016] Step 1.1.2: Randomly select two functions with the same name compiled from the same open source software or firmware but with different architectures or optimization levels as similar function pairs, and the remaining functions as dissimilar function pairs;
[0017] Step 1.1.3: All similar function pairs and dissimilar function pairs together constitute a function sample database.
[0018] Furthermore, the step 1.2 includes:
[0019] Step 1.2.1: for each function body in the function sample database, obtain its assembly program by disassembling, and obtain the control flow graph of each function body based on this;
[0020] Step 1.2.2: Analyze the assembly program and identify the library function call instructions, specific operations or operation sequences in each function body as sensitive instructions;
[0021] Step 1.2.3: For the control flow graph of each function body, take each sensitive instruction in each function as the starting point, use the reverse slicing method to extract all function instructions related to the execution of the sensitive instruction, and form a sensitive slice set corresponding to the sensitive instruction;
[0022] Step 1.2.4: Randomly select a sensitive slice as a sample slice, select qualified sensitive slices from similar functions of the function where the sample slice is located, and form a similar sensitive slice sample pair. The remaining sensitive slices are non-similar sensitive slice sample pairs; the judgment conditions of the similar sensitive slice sample pair include that the sensitive instructions of the sample slices are the same and the function call type difference is minimal.
[0023] Furthermore, the step 1.3 includes:
[0024] Step 1.3.1: Combine the control flow graph of the function where each sensitive slice is located and determine all basic blocks in the control flow graph involved in the sensitive slice;
[0025] Step 1.3.2: Analyze the statistical characteristics of each basic block. The statistical characteristics of each basic block are expressed using the vector SF=<Sum,Str,Con,Tra,Mat,Sub,Out> Indicates, where Sum represents the total number of instructions, Str represents the number of string constants, Con represents the number of numeric constants, Tra represents the number of transfer instructions, Mat represents the number of arithmetic instructions, Sub represents the number of subroutine calls, and Out represents the number of outgoing calls;
[0026] Step 1.3.3: Analyze the function call signatures in each basic block based on the instruction set and function call characteristics of different architectures. This includes: determining whether there are calls to Glibc API functions in the basic block. If so, analyze the name of the called function. Use the names and parameters of all API functions in Glibc as strings, obtain the corresponding digests of the API functions through a hash algorithm, and extract the digests of the first 10 API functions involved in the sensitive slice as the call signature. If there are less than 10 API functions involved, the call signature of the sensitive slice is set to 0.
[0027] Step 1.3.4: Based on the control flow graph of the function where the sensitive slice is located, all basic blocks involved in the sensitive slice and the original edges between these basic blocks constitute the structural characteristics of the sensitive slice.
[0028] Furthermore, the step 104 includes:
[0029] Step 1.4.1: Divide the sensitive slice sample pairs in the function sample database into two disjoint sample sets in a ratio of 4:1, one for model training and one for testing;
[0030] Step 1.4.2: Select a sensitive slice sample pair from the training set and send the two corresponding original features in the sample pair to the two sensitive slice embedding generation modules in the twin network. The two original feature vectors are first processed by Structure2vec and then sent to the fully connected layer FC for processing to obtain a final embedding E1 and E2 respectively; perform similarity calculation on E1 and E2;
[0031] Step 1.4.3: Calculate L2 Loss according to the BatchSize setting, and modify the relevant parameters in the Structure2vec and the fully connected layer based on L2 Loss through backpropagation;
[0032] Step 1.4.4: Repeat steps 1.4.2 and 1.4.3 until all samples in the training set are trained. A detection model is obtained, and the AUC and L2 loss of the model on the training set are recorded.
[0033] Step 1.4.5: Repeat steps 1.4.2 to 1.4.4 until the set upper limit of training cycles is reached, and save all detection models and their corresponding AUC and L2 Loss;
[0034] Step 1.4.6: Select one of the detection models in step 1.4.5 to configure the twin network;
[0035] Step 1.4.7: Select a sensitive slice sample pair from the test set, and feed the two original features corresponding to the sample into the sensitive slice embedding generation module in the Siamese network, and calculate the L2 loss based on the similarity results of the detection;
[0036] Step 1.4.8: Repeat step 1.4.7 until every sample in the test set has been tested;
[0037] Step 1.4.9: Calculate AUC and L2 Loss on the test set;
[0038] Step 1.4.10: Repeat steps 1.4.6 to 1.4.9 until each model in step 1.4.5 has been tested, saving the AUC and L2 Loss of each test round.
[0039] Step 1.4.11: Based on the test results of step 1.4.10, optimize the number of training cycles, batch size, and learning rate, and repeat steps 1.4.2 to 1.4.11 until the AUC and L2 loss on the test set meet the requirements. Record the optimal detection model obtained at this time as the function similarity detection model.
[0040] Furthermore, the step 2 includes:
[0041] Step 2.1: Get the target function body;
[0042] Step 2.2: Identify sensitive instructions in the target function body, and extract sensitive slices based on the sensitive instructions;
[0043] Step 2.3: Determine the original features corresponding to the sensitive slice sample;
[0044] Step 2.4: Input the original features of the target function body into the function similarity detection model to obtain the similarity detection result of the target function body.
[0045] In a second aspect, the present invention provides a function fine-grained similarity detection system based on sensitive slices, comprising:
[0046] A model building and training module is used to build a function sample database and obtain sensitive slices of each function in the function sample database, and train a function similarity detection model based on the sensitive slices; wherein the function similarity detection model is a twin network structure including two identical sensitive slice embedding generation modules, and the sensitive slice embedding generation module includes Structure2vec and a fully connected layer;
[0047] The function similarity detection module is used to perform similarity detection of the target function based on the function similarity detection model.
[0048] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the program.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program executes the method described above when executed by a processor.
[0050] The beneficial effects of the present invention are:
[0051] This paper specifically proposes a fine-grained function slice similarity detection method, transforming the feature analysis method based on the function as a whole into a discrete feature analysis process based on the function's sensitive slices. On the one hand, the granularity of function feature characterization is refined to the sensitive slice level, eliminating the interference of operations unrelated to specific sensitive instructions in the function on its features. On the other hand, the feature of a function is transformed into an element of the feature space composed of the set of sensitive slice features contained in the function, which is conducive to the detection of whether a function contains specific slice features, providing a new method for solving the problem of code inclusion detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A schematic diagram of a flow chart of a method for fine-grained function similarity detection based on sensitive slices provided by an embodiment of the present invention;
[0053] Figure 2 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0055] Existing function similarity detection methods for function vulnerability discovery have the following limitations:
[0056] (1) Using the characteristics of the vulnerable function as vulnerability characteristics actually introduces many features that may be completely irrelevant to the vulnerability (for example, a buffer overflow vulnerability, assuming that its vulnerability only involves n instructions, while the total number of instructions in the function it contains may be much higher than n), resulting in the obscuring of many characteristics of the vulnerability itself. Therefore, similar functions detected by this type of method may not have the vulnerability contained in the vulnerable function, and some functions with such vulnerabilities may not be detected.
[0057] (2) Using the entire function as the analysis object cannot solve the inclusion problem. The inclusion problem refers to the situation where two objects are not similar, but one object contains the other. Suppose a function is not similar to a vulnerable function, but it completely contains all the instructions required for a vulnerability in a vulnerable function, that is, it has the same vulnerability problem. If the function is tested based on its overall characteristics, the function will be judged as non-vulnerable due to its dissimilarity to the vulnerable function, which is obviously incorrect.
[0058] In view of the limitations of the above-mentioned prior art, the present invention takes binary functions as the analysis object. First, the sensitive operations in the function that are highly correlated with vulnerabilities and other security issues are analyzed and located; then, based on the backward slicing analysis, the function sensitive slices starting from the above-mentioned sensitive operations are obtained; finally, a slice feature model that integrates statistical features, structural features and behavioral features is constructed to characterize the sensitive slice features, and the features of each sensitive slice will be used as the effective features of its corresponding function. The present invention transforms the feature analysis method with the function as a whole as the object into a discrete feature analysis process with the function sensitive slice as the object. On the one hand, the granularity of the function feature characterization is refined to the sensitive slice level, eliminating the interference of operations in the function that are not related to specific sensitive instructions on its features; on the other hand, the feature of a function is transformed into an element in the feature space composed of the set of all sensitive slice features contained in the function, which is conducive to the detection of whether a function contains specific slice features, and provides a new method for solving the code inclusion detection problem.
[0059] Example 1
[0060] like Figure 1 As shown, an embodiment of the present invention provides a method for fine-grained function similarity detection based on sensitive slices, including:
[0061] Step 1: Build a function sample database and obtain the sensitive slices of each function in the function sample database. Based on the sensitive slices, train a function similarity detection model. The function similarity detection model is a twin network structure containing two identical sensitive slice embedding generation modules. The sensitive slice embedding generation module contains Structure2vec and a fully connected layer. It includes:
[0062] S101: Build a function sample database based on open source software or firmware;
[0063] (1) Collect open source software or firmware datasets, build a cross-compilation environment, and generate binary programs for different architectures, such as ARM, MIPS, and PowerPC. Different optimization levels are set during compilation, and multiple binary programs with different compilation optimization levels can be generated for each architecture. In this embodiment, optimization levels 00 to 03 are set, corresponding to the generation of four binary programs with four different compilation environment optimization levels.
[0064] (2) Randomly select two functions with the same function name under different architectures or different optimization levels compiled from the same open source software or firmware as similar function pairs, and the rest of the functions are non-similar function pairs.
[0065] (3) All similar function pairs and dissimilar function pairs together constitute the function sample database.
[0066] S102: Identify sensitive instructions of the function in the function sample database, and extract sensitive slices corresponding to the function in the function sample database based on the sensitive instructions; the sensitive instructions include library function call instructions, specific operations or operation sequences; the specific operations or operation sequences are specific operations or operation sequences with known vulnerabilities.
[0067] (1) For each function body in the function sample database, its assembly program is obtained by disassembly, and on this basis, the control flow graph of each function body is obtained.
[0068] (2) Analyze the assembly program and identify the call instructions, specific operations or operation sequences of library functions such as Glibc in each function body as sensitive instructions.
[0069] Identifying the call instructions of library functions such as Glibc in each function body ensures that sensitive instruction slices have a strong ability to express the functional characteristics of the function. On the other hand, due to the close correlation between the call instructions of library functions such as Glibc and device firmware vulnerabilities, it can enhance the sensitive instruction slice's ability to characterize typical vulnerabilities and better adapt to vulnerability feature detection. Finally, due to the widespread application of library functions such as Glibc, using such instructions as sensitive instructions also has strong adaptability. In addition, only using library function call instructions such as Glibc as sensitive instructions, the corresponding function slices cannot fully cover the specific operations or operation sequences of certain known vulnerabilities. Therefore, inspired by technologies such as malicious code scanning and analysis, by artificially designating certain specific operations or operation sequences as sensitive operations, we can fully characterize all aspects of the function characteristics, further improve detection capabilities, and enhance the scalability of the method.
[0070] (3) For the control flow graph of each function body, taking each sensitive instruction in each function as the starting point, the reverse slicing method is used to extract all function instructions related to the execution of the sensitive instruction to form a sensitive slice set corresponding to the sensitive instruction.
[0071] (4) Randomly select a sensitive slice as a sample slice, select sensitive slices that meet the conditions from the similar functions of the function where the sample slice is located, and form a similar sensitive slice sample pair. The remaining sensitive slices are non-similar sensitive slice sample pairs. The judgment conditions for similar sensitive slice sample pairs include that the sensitive instructions of the sample slices are the same and the function call type difference is minimal. Among them, similar sensitive slice pairs are marked as 1, and non-similar sensitive slice pairs are marked as 0.
[0072] S103: Determine original features corresponding to sensitive slices of the function in the function sample database; wherein the original features include basic block statistical features, call features, and structural features.
[0073] (1) Combine the control flow graph of the function where each sensitive slice is located and determine all basic blocks in the control flow graph involved in the sensitive slice;
[0074] (2) Analyze the statistical characteristics of each basic block. The statistical characteristics of each basic block are expressed using the vector SF =<Sum,Str,Con,Tra,Mat,Sub,Out> Indicates, where Sum represents the total number of instructions, Str represents the number of string constants, Con represents the number of numeric constants, Tra represents the number of transfer instructions, Mat represents the number of arithmetic instructions, Sub represents the number of subroutine calls, and Out represents the number of outgoing calls;
[0075] (3) According to the instruction set and function call characteristics under different architectures, the function call characteristics in each basic block are analyzed, including: determining whether there is a call to the Glibc API function in the basic block. If so, the name of the called function is analyzed. The names and parameters of all API functions in Glibc are used as strings, and the corresponding summary of the API function is obtained through the hash algorithm. The corresponding summaries of the first 10 API functions involved in the sensitive slice are intercepted as the call characteristics. If there are less than 10 API functions, the call characteristics of the sensitive slice are 0. Therefore, the function call feature vector of each sensitive slice is a 10-dimensional vector, and each dimension corresponds to the hash value of the API or is 0.
[0076] (4) Based on the control flow graph of the function where the sensitive slice is located, all the basic blocks involved in the sensitive slice and the original edges between these basic blocks constitute the structural characteristics of the sensitive slice.
[0077] S104: Training a function similarity detection model based on original features of functions in the function sample database.
[0078] (1) The sensitive slice sample pairs in the function sample database are divided into two non-overlapping sample sets in a ratio of 4:1, which are used for model training and testing respectively.
[0079] (2) A sensitive slice sample pair is selected from the training set, and the two corresponding original features in the sample pair are respectively sent to the two sensitive slice embedding generation modules in the twin network. Among them, the two original feature vectors are first processed by Structure2vec and then sent to the fully connected layer FC for processing to obtain a final embedding E1 and E2 respectively; the similarity of E1 and E2 is calculated, and the cosine distance is used as the metric to measure the similarity of the two sensitive slices.
[0080] (3) Calculate L2 Loss according to the BatchSize setting, and modify the relevant parameters in Structure2vec and the fully connected layer based on L2 Loss through back propagation.
[0081] (4) Repeat steps 1.4.2 and 1.4.3 until all samples in the training set are trained, and a detection model is obtained. The AUC and L2 Loss of the model on the training set are recorded.
[0082] (5) Repeat steps 1.4.2 to 1.4.4 until the set upper limit of training cycles is reached, and save all detection models and their corresponding AUC and L2 Loss.
[0083] (6) Select one of the detection models in step 1.4.5 to configure the twin network.
[0084] (7) Select a sensitive slice sample pair from the test set, and send the two original features corresponding to the sample into the sensitive slice embedding generation module in the twin network respectively, and calculate the L2 Loss based on the similarity results of the detection.
[0085] (8) Repeat step 1.4.7 until every sample in the test set has been tested.
[0086] (9) Calculate the AUC and L2 Loss on the test set.
[0087] (10) Repeat steps 1.4.6 to 1.4.9 until each model in step 1.4.5 has been tested, and save the AUC and L2 Loss of each round of testing.
[0088] (11) Based on the test results of step 1.4.10, optimize the number of training cycles, BatchSize, and learning rate, and repeat steps 1.4.2 to 1.4.11 until the AUC and L2 Loss on the test set meet the requirements, and record the optimal detection model obtained at this time as the function similarity detection model.
[0089] Step 2: Perform similarity detection on the target function based on the function similarity detection model. This involves detecting whether a certain software or firmware has a function that is identical or similar to the function of interest at the sensitive slice granularity, including:
[0090] S201: Obtain a target function body from an object of interest (eg, malicious code, vulnerability code, or copyright protection code).
[0091] S202: Identify sensitive instructions in the target function body and extract sensitive slices based on the sensitive instructions. The specific process is the same as S102.
[0092] S203: Determine the original features corresponding to the sensitive slice samples. The specific process is the same as S103.
[0093] S204: Inputting the original features of the target function body into the function similarity detection model to obtain a similarity detection result of the target function body.
[0094] (1) The original features of all sensitive slices of each function in the target function dataset are fed into the sensitive slice embedding generation module in the twin network to obtain the corresponding embedding of the feature. The embeddings of all sensitive slices are stored in the target function sensitive slice embedding database.
[0095] (2) Based on the different situations of the software (firmware) to be tested, decompression, structural analysis, disassembly and other processing are carried out to extract all the function objects contained therein.
[0096] (3) For each function in the software (firmware) to be detected, refer to the operations from S202 to S204 to extract the embeddings of all sensitive slices in the function, and store them in the sensitive slice embedding database of the function to be detected.
[0097] (4) For each target sensitive slice embedding in the target function sensitive slice embedding database, calculate the similarity between the embedding and each embedding in the target function sensitive slice embedding database, and output the similarity ranking.
[0098] Based on the obtained similarity scores and rankings, it is possible to quickly locate whether there are functions in the software (firmware) to be tested that are identical or similar to the target function of interest at the sensitive slice granularity. This has important reference value for further analysis to determine whether there are vulnerabilities, malicious code, or copyright infringement problems in the software (firmware) to be tested.
[0099] Example 2
[0100] In combination with the methods disclosed in the above embodiments, this embodiment describes the application of the present invention in vulnerability detection for the CVE-2016-2824 vulnerability. The CVE-2016-2824 vulnerability exists in the doapr_outch function of OpenSSL, and the affected versions include 1.0.1f and others.
[0101] (1) According to the vulnerability information of CVE-2016-2842, the openssl1.0.1f software was obtained, and the vulnerable doapr_outch function was extracted from its binary version with ARM architecture and compilation optimization level O2, and this function was used as the target vulnerable function.
[0102] (2) Identify the sensitive instructions in the doapr_outch function, extract the sensitive slices in the doapr_outch function, and construct a target sensitive slice set. The specific process is the same as S102.
[0103] (3) Extracting the original features of all sensitive slices of the doapr_outch function. The specific process is the same as S103.
[0104] (4) Obtain the embedded features of all sensitive slices of the doapr_outch function. The specific process is the same as S204.
[0105] (5) In order to illustrate the detection process and effect, based on the openssl1.0.1f source code, three openssl1.0.1f binary programs corresponding to the three optimization levels of O0, O1, and O3 under the ARM architecture are generated through different compilation environments and parameter settings; then, these three binary programs are used as the programs to be tested to detect whether there are sensitive slices similar to the target vulnerability function doapr_outch.
[0106] (6) For each binary program, refer to the processing steps S201-S204 to obtain the similarity rankings between all sensitive slices of the function to be detected and the sensitive slices in the target function doapr_outch.
[0107] Through actual detection, the results shown in the following table are obtained. The table lists the important sensitive slice information in the vulnerable function doapr_outch, including the number of sensitive slices and the library functions called by the corresponding sensitive instructions. By analyzing the functions of each function and the CVE description of the vulnerability, it can be found that these sensitive slices are closely related to the implementation of the core functions of the function and the existence of the vulnerability. Among all the similarity test results, all sensitive slices are ranked according to their similarity (arranged in descending order of similarity). The table lists the similarity values and rankings of the sensitive slices that are truly similar to the sample slices in the test results, all of which achieve the Top2 hit test level, verifying the high accuracy of fine-grained similarity detection of sensitive slices.
[0108]
[0109] In addition, since the trigger of the CVE-2016-2842 vulnerability is a memory overflow caused by calling memcpy, the memcpy instruction sensitive slice has a more important indicator significance. If the sensitive slice in a function is highly similar to the memcpy sensitive slice, it may mean that the function also has a similar vulnerability problem with memcpy as the sensitive instruction, but the function as a whole may be significantly different from the doapr_outch function. In other words, the present invention can effectively address and solve the inclusion class detection problem.
[0110] Example 3:
[0111] Based on the above embodiment, this embodiment further provides a function fine-grained similarity detection system based on sensitive slices, including:
[0112] The model building and training module is used to build a function sample database and obtain the sensitive slices of each function in the function sample database, and train a function similarity detection model based on the sensitive slices. The function similarity detection model is a twin network structure containing two identical sensitive slice embedding generation modules. The sensitive slice embedding generation module includes Structure2vec and a fully connected layer.
[0113] The function similarity detection module is used to perform similarity detection of the target function based on the function similarity detection model.
[0114] Example 4:
[0115] Based on the above embodiments, Figure 2As shown, this embodiment further provides an electronic device, which may include: a processor (processor) 201, a communication interface (Communications Interface) 202, a memory (memory) 203 and a communication bus 204, wherein the processor 201, the communication interface 202, and the memory 203 communicate with each other via the communication bus 204. The processor 201 may call the logic instructions in the memory 203 to execute the method provided in the above embodiment, for example, including:
[0116] Step 1: Build a function sample database and obtain sensitive slices for each function in the database. Train a function similarity detection model based on the sensitive slices. The function similarity detection model is a twin network structure containing two identical sensitive slice embedding generation modules, each containing Structure2vec and a fully connected layer. Step 2: Perform similarity detection on the target function based on the function similarity detection model.
[0117] In addition, when the logic instructions in the above-mentioned memory 203 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0118] Example 5:
[0119] Based on the above embodiments, this embodiment further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method provided in the above embodiments is implemented, for example, including:
[0120] Step 1: Build a function sample database and obtain sensitive slices for each function in the database. Train a function similarity detection model based on the sensitive slices. The function similarity detection model is a twin network structure containing two identical sensitive slice embedding generation modules, each containing Structure2vec and a fully connected layer. Step 2: Perform similarity detection on the target function based on the function similarity detection model.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for fine-grained function similarity detection based on sensitive slicing, characterized in that: include: Step 1: Build a function sample database and obtain the sensitive slice of each function in the function sample database, and train a function similarity detection model based on the sensitive slice; The function similarity detection model is a twin network structure comprising two identical sensitive slice embedding generation modules, wherein the sensitive slice embedding generation module comprises Structure2vec and a fully connected layer; Step 2: Perform similarity detection on the target function based on the function similarity detection model.
2. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 1, characterized in that: The step 1 comprises: Step 1.1: Build a function sample database based on open source software or firmware; Step 1.2: Identify sensitive instructions of the functions in the function sample database, and extract sensitive slices corresponding to the functions in the function sample database based on the sensitive instructions; the sensitive instructions include library function call instructions, specific operations or operation sequences; the specific operations or operation sequences are specific operations or operation sequences with known vulnerabilities; Step 1.3: determining original features corresponding to sensitive slices of the function in the function sample database; wherein the original features include basic block statistical features, call features, and structural features; Step 1.4: training the function similarity detection model based on the original features of the functions in the function sample database.
3. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 2, characterized in that: The step 1.1 includes: Step 1.1.1: Collect open source software or firmware datasets, generate binary programs for different architectures through cross-compilation, and set different optimization levels during compilation. For each architecture, multiple binary programs with different compilation optimization levels can be generated. Step 1.1.2: Randomly select two functions with the same name compiled from the same open source software or firmware but with different architectures or optimization levels as similar function pairs, and the remaining functions as dissimilar function pairs; Step 1.1.3: All similar function pairs and dissimilar function pairs together constitute a function sample database.
4. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 2, characterized in that: The step 1.2 includes: Step 1.2.1: for each function body in the function sample database, obtain its assembly program by disassembling, and obtain the control flow graph of each function body based on this; Step 1.2.2: Analyze the assembly program and identify the library function call instructions, specific operations or operation sequences in each function body as sensitive instructions; Step 1.2.3: For the control flow graph of each function body, take each sensitive instruction in each function as the starting point, use the reverse slicing method to extract all function instructions related to the execution of the sensitive instruction, and form a sensitive slice set corresponding to the sensitive instruction; Step 1.2.4: Randomly select a sensitive slice as a sample slice, select qualified sensitive slices from similar functions of the function where the sample slice is located, and form a similar sensitive slice sample pair. The remaining sensitive slices are non-similar sensitive slice sample pairs; the judgment conditions of the similar sensitive slice sample pair include that the sensitive instructions of the sample slices are the same and the function call type difference is minimal.
5. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 2, characterized in that: The step 1.3 includes: Step 1.3.1: Combine the control flow graph of the function where each sensitive slice is located and determine all basic blocks in the control flow graph involved in the sensitive slice; Step 1.3.2: Analyze the statistical characteristics of each basic block. The statistical characteristics of each basic block are expressed using the vector SF=<Sum,Str,Con,Tra,Mat,Sub,Out> Indicates, where Sum represents the total number of instructions, Str represents the number of string constants, Con represents the number of numeric constants, Tra represents the number of transfer instructions, Mat represents the number of arithmetic instructions, Sub represents the number of subroutine calls, and Out represents the number of outgoing calls; Step 1.3.3: Analyze the function call signatures in each basic block based on the instruction set and function call characteristics of different architectures. This includes: determining whether there are calls to Glibc API functions in the basic block. If so, analyze the name of the called function. Use the names and parameters of all API functions in Glibc as strings, obtain the corresponding digests of the API functions through a hash algorithm, and extract the digests of the first 10 API functions involved in the sensitive slice as the call signature. If there are less than 10 API functions involved, the call signature of the sensitive slice is set to 0. Step 1.3.4: Based on the control flow graph of the function where the sensitive slice is located, all basic blocks involved in the sensitive slice and the original edges between these basic blocks constitute the structural characteristics of the sensitive slice.
6. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 2, characterized in that: The step 1.4 includes: Step 1.4.1: Divide the sensitive slice sample pairs in the function sample database into two disjoint sample sets in a ratio of 4:1, one for model training and one for testing; Step 1.4.2: Select a sensitive slice sample pair from the training set and send the two corresponding original features in the sample pair to the two sensitive slice embedding generation modules in the twin network. The two original feature vectors are first processed by Structure2vec and then sent to the fully connected layer FC for processing to obtain a final embedding E1 and E2 respectively; perform similarity calculation on E1 and E2; Step 1.4.3: Calculate L2 Loss according to the BatchSize setting, and modify the relevant parameters in the Structure2vec and the fully connected layer based on L2 Loss through backpropagation; Step 1.4.4: Repeat steps 1.4.2 and 1.4.3 until all samples in the training set are trained. A detection model is obtained, and the AUC and L2 loss of the model on the training set are recorded. Step 1.4.5: Repeat steps 1.4.2 to 1.4.4 until the set upper limit of training cycles is reached, and save all detection models and their corresponding AUC and L2 Loss; Step 1.4.6: Select one of the detection models in step 1.4.5 to configure the twin network; Step 1.4.7: Select a sensitive slice sample pair from the test set, and feed the two original features corresponding to the sample into the sensitive slice embedding generation module in the Siamese network, and calculate the L2 loss based on the similarity results of the detection; Step 1.4.8: Repeat step 1.4.7 until every sample in the test set has been tested; Step 1.4.9: Calculate AUC and L2 Loss on the test set; Step 1.4.10: Repeat steps 1.4.6 to 1.4.9 until each model in step 1.4.5 has been tested, saving the AUC and L2 Loss of each test round. Step 1.4.11: Based on the test results of step 1.4.10, optimize the number of training cycles, batch size, and learning rate, and repeat steps 1.4.2 to 1.4.11 until the AUC and L2 loss on the test set meet the requirements. Record the optimal detection model obtained at this time as the function similarity detection model.
7. The method for detecting fine-grained similarity of functions based on sensitive slices according to claim 1, characterized in that: The step 2 includes: Step 2.1: Get the target function body; Step 2.2: Identify sensitive instructions in the target function body, and extract sensitive slices based on the sensitive instructions; Step 2.3: Determine the original features corresponding to the sensitive slice sample; Step 2.4: Input the original features of the target function body into the function similarity detection model to obtain the similarity detection result of the target function body.
8. A function fine-grained similarity detection system based on sensitive slicing, characterized by: include: A model building and training module is used to build a function sample database and obtain sensitive slices of each function in the function sample database, and train a function similarity detection model based on the sensitive slices; The function similarity detection model is a twin network structure comprising two identical sensitive slice embedding generation modules, wherein the sensitive slice embedding generation module comprises Structure2vec and a fully connected layer; The function similarity detection module is used to perform similarity detection of the target function based on the function similarity detection model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.