Method for Extracting and Identifying Compiler Fusion Features Based on Disassembly
By extracting the compiler's statistical and correlation features based on disassembly, and combining chi-square test and machine learning model, the problem of low accuracy of compiler optimization level recognition is solved, and efficient recognition of the compiler family, version and optimization level is achieved.
Patent Information
- Application Number
- CN202310746488.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-21
AI Technical Summary
The prior art recognizes optimization levels with relatively small differences in compilers, and the recognition accuracy is not high, especially the recognition accuracy of optimization levels such as O0, O1 and O2 is insufficient, and deep learning methods fail to fully reflect the characteristics of the compiler.
The compiler's statistical and associated features are extracted by a disassembly-based method, and representative features are selected in combination with the chi-square test feature selection method, and four machine learning models, SVM, LightGBM, XGBoost and RF, are used to identify the compiler family, version and optimization level.
The recognition accuracy of the compiler family, version and optimization level has been improved, especially the recognition effect of optimization levels such as O0, O1 and O2 has been significantly improved, achieving higher classification accuracy.
Smart Images

Figure CN116720071B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software security, and specifically to a method for extracting and identifying compiler fusion features based on disassembly. Background Art
[0002] In recent years, software supply chain attack incidents have occurred frequently, which have seriously threatened the security of the software industry. There are mainly five types of attacks on the software supply chain: development tool contamination, manufacturer backdoors, source code contamination, bundled downloads, and upgrade hijacking. Among the five types of attacks, attacking the development tool compiler has lower costs, a wider scope of influence, and higher efficiency, which has attracted wide attention from attackers, but this is the attack behavior that security analysts are most likely to overlook.
[0003] In the XcodeGhost compiler attack incident in 2015, attackers attracted software developers to download an unofficial Xcode compiler implanted with malicious code. This contaminated Xcode compiler was widely spread within half a year, and a large number of domestic APPs compiled by this compiler were implanted with malicious code. These APPs could upload user privacy to a specific URL address, and more than 100 million users were affected in this incident.
[0004] In 1984, K.T. Hopkins proposed that attackers could contaminate all compiled and released software by attacking the compiler. Compared with attacking the software itself, attacking the software development tool has lower difficulty and cost, but a wider scope of influence. As an essential tool in software development, if the compiler is contaminated, then all software compiled by this compiler will be contaminated. And the compiler is at the upstream of the software supply chain, and subsequent multiple transmissions on the supply chain will cover up the attack behavior of the compiler, making it difficult for security personnel to prevent and handle. Compiler attacks have attracted the attention of attackers, and the number and influence of such attack incidents will further expand in the future, seriously affecting social stability and national security.
[0005] Different configuration information of the compiler will affect the representation of binary files, thereby affecting binary analysis tasks. Therefore, compiler identification is an important part of binary toolchain analysis, and it has important contribution value in reverse engineering, malware analysis, and identifying untrustworthy software components in the software supply chain. The task of security investigators is usually to analyze and attribute binary files obtained from malicious activities and quickly and reliably check these files. These binary files usually contain intelligence about the tactics, techniques, and procedures of the opponent. Reverse-identifying the compiler family, version, and optimization information from binary files can accelerate the analysis. Therefore, it is very meaningful to develop an efficient and automated method for compiler identification from binary files.
[0006] Existing research mainly focuses on identifying compilers based on deep learning. This method generally only performs simple processing on binary files and then uses the network to autonomously learn the internal features of binary files. It cannot fully reflect the features of compilers. Moreover, the features extracted from binary files by this method contain too many redundant features, resulting in low recognition accuracy for compiler optimization levels with small differences. Therefore, existing research generally only identifies two optimization levels, O0 and O2, or combines the two optimization levels, O2 and O3, into a new optimization level OH, and then identifies three optimization levels, O0, O1, and OH. Summary of the Invention
[0007] The present invention proposes a method for extracting and identifying compiler fusion features based on disassembly, which solves the problem of low recognition accuracy for compiler optimization levels with small differences. A C-language source code dataset is generated, and more representative compiler features are extracted based on executable files and combined with machine learning to identify the family, version, and optimization level of compilers.
[0008] The method for extracting and identifying compiler fusion features based on disassembly of the present invention is specifically implemented according to the following steps:
[0009] Step 1, dataset construction;
[0010] Step 2, data preprocessing: Use the IDA Pro disassembly tool to disassemble the executable file into a disassembly file, that is, an asm file;
[0011] Step 3, feature extraction: Extract statistical features and correlation features from the disassembly file obtained in Step 2. The statistical features include register usage frequency and opcode usage frequency;
[0012] Step 4, feature preprocessing: Complete dataset partitioning for the statistical features and correlation features extracted in Step 3 to obtain 8 datasets. Use the chi-square test feature selection method to screen the extracted statistical features and correlation features respectively, select the feature set with the top 40% chi-square scores, and fuse the screened statistical features and correlation features as the basis for subsequent compiler recognition;
[0013] Step 5, compiler recognition: Use four machine learning models, SVM, LightGBM, XGBoost, and RF, to conduct classification experiments on the dimensionality-reduced statistical features, correlation features, and fusion features on the 8 datasets in Step 4 to achieve the recognition of the compiler family, version, and optimization level;
[0014] The specific implementation of the said Step 1 is as follows:
[0015] Step 1.1, Download the CSmith test tool in the Linux operating system, configure the CSmith environment, and generate 10,000 C language source codes as a dataset, denoted as Dataset1 dataset;
[0016] Step 1.2, In order to evaluate the extracted features from three perspectives: the family, version, and optimization level of the compiler, with the optimization levels being O0, O1, O2, and O3, download three types of compilers, namely gcc8.1.0, gcc9.2.0, and clang10.0.0, in the Win10 operating system;
[0017] Step 1.3, Write a python batch script to compile the 10,000 C language source code files in the Dataset1 dataset in Step 1.1 into corresponding executable code files using the four optimization levels of the three compilers in Step 1.2, with the optimization levels being O0, O1, O2, and O3; 10,000 executable files are generated for each optimization level, and finally 120,000 executable files are obtained; Use the 120,000 executable files as the subsequent experimental dataset, denoted as Dataset2 dataset. The category and quantity description of the Dataset2 dataset is shown in Table 1;
[0018] Table 1 Description of Dataset2 dataset
[0019]
[0020]
[0021] The specific implementation of Step 2 is as follows:
[0022] Preprocess the executable files in the Dataset2 dataset; Write an IDA Python automation script analysis program to call the IDA Pro tool to decompile the Dataset2 dataset in Step 1.3 into disassembly files with the suffix.asm in batch, and denote the disassembly dataset as Data. Then extract statistical features and correlation features based on Data;
[0023] The specific implementation of Step 3 is as follows:
[0024] For the convenience of subsequent processing, fuse the features according to the labels (for example, combine two features with the label gcc8.1.0_O0 / / 1 into a one-dimensional feature vector) and store them in a csv file; In the csv file, the first n columns are statistical features, and the remaining columns are correlation features;
[0025] Step 3.1, Extract the register usage frequency from the disassembly files obtained in Step 2;
[0026] Step 3.2, extract the opcode usage frequency from the disassembly file obtained in Step 2;
[0027] Step 3.3, collectively refer to the register usage frequency extracted in Step 3.1 and the opcode usage frequency extracted in Step 3.2 as statistical features. Considering that the statistical features lack the local information of the assembly code, an N-gram feature extraction method based on the statistical feature sequence is adopted to obtain correlation features, where N = 2, and N-gram refers to a continuous subsequence composed of N feature units;
[0028] The specific implementation of Step 4 is as follows:
[0029] Step 4.1, re-partitioning of the feature set: The statistical features and correlation features extracted from the Dataset2 dataset are fused using the merge function in the panda library and stored in a CSV file. The statistical features and correlation features are divided into 7 sub-datasets according to the compiler's family, version, and optimization level. The descriptions of the 7 sub-datasets are shown in Table 2. Finally, 8 datasets are generated, and the 8 datasets are: Dataset2 and 7 sub-datasets;
[0030] Step 4.2, because there are a large number of redundant features in the statistical features and correlation features extracted in Step 3.3, in order to improve the accuracy and training speed of model training, the chi-square test feature selection method is used to select the extracted statistical features and correlation features on the 8 datasets respectively, and the feature sets with the top 40% chi-square scores are selected. The fused features after dimensionality reduction are used as the basis for identifying the compiler later. The statistical features are described below;
[0031] Step 4.3, feature fusion: The statistical features and correlation features with the same label are fused. Finally, a total of 972-dimensional features are extracted for each asm file and stored in a CSV file; the first 134 dimensions are statistical features (the first 42-dimensional features are register features, and the last 92-dimensional features are opcode features), and the remaining 838 columns are correlation features; then, the extracted statistical features and correlation features are uniformly divided into training sets and test sets, and then the chi-square test feature selection method is used to select each type of feature according to the columns where the features are located, and the selected features are fused.
[0032] Table 2 Dataset description
[0033]
[0034] Preferably, the specific content of Step 3.1 is as follows:
[0035] (1) Count the usage frequencies of 42 commonly used registers, and take the usage frequencies of the 42 registers as a feature of the compiler. The following are the register types counted this time: eax, ebx, cs, ds, ss, es, fs, gs, ah, al, ax, bh, bl, bx, ch, cl, cx, dh, dl, dx, ecx, edx, esp, ebp, esi, edi, rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, r8, r9, r10, r11, r12, r13, r14, r15;
[0036] (2) Count the frequencies of the above registers appearing in each asm file, and extract 42-dimensional register features from each asm file.
[0037] Preferably, step 3.2 is specifically: count the usage frequencies of 92 commonly used opcodes in each asm file.
[0038] (1) Count 92 commonly used opcodes as follows: add, bt, call, cdq, cld, cli, cmc, cmp, const, cwd, daa, db, dd, dec, dw, endp, ends, faddp, fchs, fdiv, fdivp, fdivr, fild, fistp, fld, fstcw, fstcwimul, fstp, fword, fxch, imul, in, inc, ins, int, jb, je, jg, jge, jl, jmp, jnb, jno, jnz, jo, jz, lea, loope, mov, movzx, mul, near, neg, not, or, out, outs, pop, popf, proc, push, pushf, rcl, rcr, rdtsc, rep, ret, retn, rol, ror, sal, sar, sbb, scas, setb, setle, setnle, setnz, setz, shl, shld, shr, sidt, stc, std, sti, stos, sub, test, wait, xchg, xor;
[0039] (2) Count the usage frequencies of the above opcodes in each asm file.
[0040] Preferably, the N-gram feature extraction method based on the statistical feature sequence in step 3.3 is specifically:
[0041] (1) Read the disassembly files in the Data dataset line by line, and extract the registers and opcodes that appear in the statistical feature sequence in the files;
[0042] (2) Store the registers and opcodes obtained from the disassembly file in step (1) into a list. Take every two elements in the list as a tuple, and count the frequency of all such tuples in the entire list in the Data dataset. Use the tuple as the key of a dictionary and the frequency of the tuple's occurrence as the value of the dictionary. The key composed of the tuples is the N-gram feature based on statistical characteristics extracted from a disassembly file.
[0043] (3) Screen the N-gram features of the entire dataset, select the features with a frequency greater than 2500, and finally 838-dimensional N-gram features are extracted from each disassembly file.
[0044] Preferably, the specific descriptions of the 7 sub-datasets in step 4.1 are as follows:
[0045] (1) D gcc_clang The D dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2, gcc_9.2.0_O3, clang_10.0.0_O0, clang_10.0.0_O1, clang_10.0.0_O2, and clang_10.0.0_O3. These 8 types of data are divided into 2 categories according to different compiler families, with labels {gcc, clang}. This dataset is used to identify two different compiler families, gcc and clang.
[0046] (2) D gcc_ver The D dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_8.1.0_O0, gcc_8.1.0_O1, gcc_8.1.0_O2, gcc_8.1.0_O3, gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2, and gcc_9.2.0_O3. The dataset is divided into 2 categories according to different versions of the gcc compiler, with labels {gcc_8.1.0, gcc_9.2.0}. This dataset is used to identify two different versions of the gcc compiler, gcc_8.1.0 and gcc_9.2.0.
[0047] (3) D gcc_O0_O2 The D dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_9.2.0_O0 and gcc_9.2.0_O2. The dataset is divided into 2 categories according to different optimization levels, with labels {gcc_9.2.0_O0, gcc_9.2.0_O2}. This dataset is used to identify two optimization levels, O0 and O2, of the gcc 9.2.0 version compiler.
[0048] (4)D clang_O0_O2 The dataset is generated by mixing the datasets in Dataset2 with data categories of clang_10.0.0_O0 and clang_10.0.0_O2. The dataset is divided into 2 categories according to different optimization levels, with labels {clang_10.0.0_O0, clang_10.0.0_O2}. This dataset identifies the O0 and O2 optimization levels of the clang10.0.0 version compiler.
[0049] (5)D multi_O0_O2 The dataset is generated by mixing D gcc_O0_O2 and D clang_O0_O2 Two datasets are generated by mixing. The dataset is divided into 4 categories according to different optimization levels, with labels {gcc_9.2.0_O0, gcc_9.2.0_O2, clang_10.0.0_O0, clang_10.0.0_O2}. This dataset identifies 4 types under the O0 and O2 optimization levels of the gcc9.2.0 and clang10.0.0 version compilers.
[0050] (6)D gcc_O0-O3 The dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_8.1.0_O0, gcc_8.1.0_O1, gcc_8.1.0_O2, gcc_8.1.0_O3, gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2 and gcc_9.2.0_O3. The dataset is divided into 4 categories according to different optimization levels, with labels {O0, O1, O2, O3}. This dataset identifies the O0, O1, O2 and O3 optimization levels of the gcc compiler.
[0051] (7)D clang_O0-O3 The dataset is generated by mixing the datasets in Dataset2 with data categories of clang_10.0.0_O0, clang_10.0.0_O1, clang_10.0.0_O2 and clang_10.0.0_O3. The dataset is divided into 4 categories according to different optimization levels, with labels {O0, O1, O2, O3}. This dataset identifies the O0, O1, O2 and O3 optimization levels of the clang compiler.
[0052] Preferably, the chi-square test feature selection method in step 4.2 is specifically as follows:
[0053] Assume that the sample space is represented by matrix X m×n =(x1,x2,…,x n, y), where m represents the total number of samples, n represents the number of features extracted from each sample, x i represents the i-th feature, and y represents the sample category. There are 8 datasets with different numbers of samples. Therefore, here m is used to represent the number of samples in the dataset, n represents the feature category, and n is 134;
[0054] (1) Assume that each feature is independent of the label y;
[0055] (2) Use formula (3-1) to calculate the chi-square value χ i between each feature x 2 (x i , y);
[0056] (3) Sort the n chi-square values χ 2 (x1, y), χ 2 (x2, y), …, χ 2 (x n , y) from largest to smallest, and select features according to the set threshold.
[0057]
[0058] In the formula, χ 2 (x, y) represents the chi-square value calculated for each feature, x i represents the actual value of the i-th data; E i represents the expected value of the i-th data; n represents the total number of observed data;
[0059] (4) The process of using chi-square test for dimensionality reduction of correlated features is similar to the above. Finally, the dimensionality-reduced features are fused as the recognition basis for the subsequent compiler.
[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0061] In view of the problem that the current method mainly extracts coarse-grained compiler features based on deep learning and has a low recognition accuracy for compiler optimization levels with small differences, the present invention proposes a method for extracting compiler fusion features and identifying optimization levels based on disassembly. This method first proposes to extract the fusion features of statistical features (usage frequencies of common registers and opcodes) and correlation features from disassembly code as the features of the compiler. Considering that the extracted features contain redundant features and the extracted features are discrete features, the chi-square test feature selection method is then used to select each feature, and finally, the feature set with the top 40% chi-square scores is selected for fusion. The fused features are used as the basis for compiler recognition, and four machine learning models, namely SVM, LightGBM, XGBoost, and RF, are used to conduct classification experiments on the extracted single features and fused features respectively. The present invention uses a combination of the chi-square test feature selection method, the LightGBM machine learning model, and the fused features to have a good recognition effect on the family, version, and four optimization levels of the compiler. Brief Description of the Drawings
[0062] The technical solution of the present invention will be further described in detail below in conjunction with the drawings and embodiments.
[0063] Figure 1 is a flowchart of the method for extracting compiler fusion features and identifying optimization levels based on disassembly according to the present invention; Detailed Embodiments
[0064] The following will disclose multiple embodiments of the present invention in the form of diagrams. For the sake of clarity, many physical details will be described together in the following narrative. However, it should be understood that these physical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these physical details are unnecessary. In addition, for the purpose of simplifying the diagrams, some conventional structures and components will be shown in a simple schematic manner in the diagrams.
[0065] In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the premise that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0066] Such as Figure 1 The method for extracting and identifying compiler fusion features based on disassembly according to the present invention, as Figure 1 shown, is specifically implemented according to the following steps:
[0067] Step 1, dataset construction; it is specifically implemented according to the following steps:
[0068] Step 1.1, Download the CSmith test tool in the Linux operating system, configure the CSmith environment, and generate 10,000 C language source codes as a dataset, denoted as Dataset1.
[0069] Step 1.2, In order to evaluate the extracted features from three perspectives: the family, version, and optimization level of the compiler, with optimization levels of O0, O1, O2, and O3, download three types of compilers, namely gcc8.1.0, gcc9.2.0, and clang10.0.0, in the Win10 operating system;
[0070] Step 1.3, Write a python batch script to compile the 10,000 C language source code files in the Dataset1 dataset in Step 1.1 into corresponding executable code files using the four optimization levels of the three compilers in Step 1.2, with optimization levels of O0, O1, O2, and O3; 10,000 executable files are generated for each optimization level, and finally 120,000 executable files are obtained; use the 120,000 executable files as the subsequent experimental dataset, denoted as Dataset2. The category and quantity description of this dataset are shown in Table 1;
[0071] Table 1 Dataset2 dataset description
[0072]
[0073] Step 2, Data preprocessing: Use the IDA Pro disassembler tool to disassemble the executable file into a disassembled file, that is, an asm file; specifically, it is implemented according to the following steps:
[0074] Preprocess the executable files in the Dataset2 dataset; write an IDA Python automation script analysis program to call the IDA Pro tool to batch decompile the Dataset2 dataset in Step 1.3 into disassembled files with the suffix of.asm format, denote this disassembled dataset as Data, and then extract statistical features and correlation features based on Data;
[0075] Step 3, Feature extraction: Extract statistical features and correlation features from the disassembled files obtained in Step 2. The statistical features include register usage frequency and opcode usage frequency; specifically, it is implemented according to the following steps:
[0076] For the convenience of subsequent processing, the features are fused according to labels (for example, combine two features with the label gcc8.1.0_O0 / / 1 into a one-dimensional feature vector) and then stored in a csv file. In this file, the first n columns are statistical features, and the remaining columns are correlation features;
[0077] Step 3.1: Extract the register usage frequency from the disassembly file obtained in Step 2;
[0078] Step 3.2: Extract the opcode usage frequency from the disassembly file obtained in Step 2;
[0079] Step 3.3: Collectively refer to the register usage frequency extracted in Step 3.1 and the opcode usage frequency extracted in Step 3.2 as statistical features. Considering that the statistical features lack local information of the assembly code, an N-gram feature extraction method based on the statistical feature sequence is used to obtain correlation features, where N = 2, and N-gram refers to a continuous subsequence composed of N feature units;
[0080] Step 3.1 is specifically as follows:
[0081] (1) Count the usage frequency of 42 commonly used registers, and use the usage frequency of the 42 registers as a feature of the compiler. The following are the register types counted this time: eax, ebx, cs, ds, ss, es, fs, gs, ah, al, ax, bh, bl, bx, ch, cl, cx, dh, dl, dx, ecx, edx, esp, ebp, esi, edi, rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, r8, r9, r10, r11, r12, r13, r14, r15;
[0082] (2) Count the frequency of occurrence of the above registers in each asm file, and extract 42-dimensional register features from each asm file.
[0083] Step 3.2 is specifically as follows:
[0084] (1) Count 92 common opcodes as follows: add, bt, call, cdq, cld, cli, cmc, cmp, const, cwd, daa, db, dd, dec, dw, endp, ends, faddp, fchs, fdiv, fdivp, fdivr, fild, fistp, fld, fstcw, fstcwimul, fstp, fword, fxch, imul, in, inc, ins, int, jb, je, jg, jge, jl, jmp, jnb, jno, jnz, jo, jz, lea, loope, mov, movzx, mul, near, neg, not, or, out, outs, pop, popf, proc, push, pushf, rcl, rcr, rdtsc, rep, ret, retn, rol, ror, sal, sar, sbb, scas, setb, setle, setnle, setnz, setz, shl, shld, shr, sidt, stc, std, sti, stos, sub, test, wait, xchg, xor;
[0085] (2) Count the usage frequencies of the above opcodes in each asm file.
[0086] The N-gram feature extraction method based on the statistical feature sequence in Step 3.3 is specifically as follows:
[0087] (1) Read the disassembly files in the Data dataset line by line, and extract the registers and opcodes that appear in the statistical feature sequence in the files;
[0088] (2) Store the registers and opcodes obtained from the disassembly files in step (1) into a list. Take every two elements in the list as a tuple, and count the frequencies of all bigrams in the entire list in the Data dataset; Use the tuple as the key of the dictionary and the frequency of the tuple's occurrence as the value of the dictionary. The key composed of the bigram is the N-gram feature based on the statistical features extracted from a disassembly file.
[0089] (3) Screen the N-gram features of the entire dataset, select the features with a frequency greater than 2500, and finally extract 838-dimensional N-gram features for each disassembly file.
[0090] Step 4, Feature preprocessing: Divide the statistical features and correlation features extracted in Step 3 to obtain 8 datasets. Use the chi-square test feature selection method to screen the extracted statistical features and correlation features respectively, and select the feature set with the top 40% chi-square scores. The screened statistical features and correlation features are fused as the recognition basis for the subsequent compiler; specifically, it is implemented according to the following steps:
[0091] Step 4.1, Feature set re-division: The statistical features and correlation features extracted from the Dataset2 dataset are fused using the merge function in the panda library and stored in a CSV file. The statistical features and correlation features are divided into 7 sub-datasets according to the family, version, and optimization level of the compiler. The descriptions of the 7 sub-datasets are shown in Table 2. Finally, 8 datasets are generated, namely Dataset2 and 7 sub-datasets;
[0092] Step 4.2, Since there are a large number of redundant features in the statistical features and correlation features extracted in Step 3.3, in order to improve the accuracy and training speed of model training, use the chi-square test feature selection method to select the extracted statistical features and correlation features respectively on 8 datasets, and select the feature set with the top 40% chi-square scores. The dimension-reduced features are fused as the recognition basis for the subsequent compiler. The following is an explanation of the statistical features;
[0093] Step 4.3, Feature fusion: The statistical features and correlation features with the same labels are fused. Finally, a total of 972-dimensional features are extracted for each asm file and stored in a CSV file; the first 134 dimensions are statistical features (the first 42-dimensional features are register features, and the last 92-dimensional features are opcode features), and the remaining 838 columns are correlation features; then, the extracted statistical features and correlation features are uniformly divided into training sets and test sets, and then each type of feature is selected using the chi-square test feature selection method according to the columns where the features are located, and the selected features are fused.
[0094] Table 2 Dataset description
[0095]
[0096] The specific descriptions of the 7 sub-datasets in Step 4.1 are as follows:
[0097] (1) D gcc_clangThe dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2, gcc_9.2.0_O3, clang_10.0.0_O0, clang_10.0.0_O1, clang_10.0.0_O2, and clang_10.0.0_O3. These 8 types of data are divided into 2 categories according to different compiler families, with labels {gcc, clang}. This dataset is used to identify compilers of two different families, gcc and clang.
[0098] (2)D gcc_ver The dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_8.1.0_O0, gcc_8.1.0_O1, gcc_8.1.0_O2, gcc_8.1.0_O3, gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2, and gcc_9.2.0_O3. The dataset is divided into 2 categories according to different versions of the gcc compiler, with labels {gcc_8.1.0, gcc_9.2.0}. This dataset is used to identify compilers of two different versions, gcc_8.1.0 and gcc_9.2.0.
[0099] (3)D gcc_O0_O2 The dataset is generated by mixing the datasets in Dataset2 with data categories of gcc_9.2.0_O0 and gcc_9.2.0_O2. The dataset is divided into 2 categories according to different optimization levels, with labels {gcc_9.2.0_O0, gcc_9.2.0_O2}. This dataset is used to identify two optimization levels, O0 and O2, of the gcc 9.2.0 version compiler.
[0100] (4)D clang_O0_O2 The dataset is generated by mixing the datasets in Dataset2 with data categories of clang_10.0.0_O0 and clang_10.0.0_O2. The dataset is divided into 2 categories according to different optimization levels, with labels {clang_10.0.0_O0, clang_10.0.0_O2}. This dataset is used to identify two optimization levels, O0 and O2, of the clang 10.0.0 version compiler.
[0101] (5)D multi_O0_O2 The dataset is D gcc_O0_O2 and D clang_O0_O2Generated by mixing two datasets. According to different optimization levels, the dataset is divided into 4 categories, with labels {gcc_9.2.0_O0, gcc_9.2.0_O2, clang_10.0.0_O0, clang_10.0.0_O2}. This dataset identifies 4 types under the O0 and O2 optimization levels of the gcc9.2.0 and clang10.0.0 compilers.
[0102] (6)D gcc_O0-O3 The dataset is generated by mixing the datasets in Dataset2 with data categories gcc_8.1.0_O0, gcc_8.1.0_O1, gcc_8.1.0_O2, gcc_8.1.0_O3, gcc_9.2.0_O0, gcc_9.2.0_O1, gcc_9.2.0_O2, and gcc_9.2.0_O3. According to different optimization levels, the dataset is divided into 4 categories, with labels {O0, O1, O2, O3}. This dataset identifies the O0, O1, O2, and O3 optimization levels of the gcc compiler.
[0103] (7)D clang_O0-O3 The dataset is generated by mixing the datasets in Dataset2 with data categories clang_10.0.0_O0, clang_10.0.0_O1, clang_10.0.0_O2, and clang_10.0.0_O3. According to different optimization levels, the dataset is divided into 4 categories, with labels {O0, O1, O2, O3}. This dataset identifies the O0, O1, O2, and O3 optimization levels of the clang compiler.
[0104] The chi-square test feature selection method in step 4.2 is specifically as follows:
[0105] Assume that the sample space is represented by matrix X m×n =(x1,x2,…,x n ,y), where m represents the total number of samples, n represents the number of features extracted from each sample, x i represents the i-th feature, and y represents the sample category. There are 8 datasets with different sample numbers, so here m is used to represent the number of samples in the dataset, and n represents the feature category (n is 134 in this section);
[0106] (1) Assume that each feature is independent of the label y;
[0107] (2) Use formula (3-1) to calculate the chi-square value χ i of each feature x 2 (x i ,y);
[0108] (3) Sort the n χ 2 (x1, y), χ 2 (x2, y), …, χ 2 (x n , y) in descending order and select features according to the set threshold.
[0109]
[0110] In the formula, χ 2 (x, y) represents the chi-square value calculated for each feature, x i represents the actual value of the i-th data; E i represents the expected value of the i-th data; n represents the total number of observed data;
[0111] (4) The process of dimensionality reduction for associated features using chi-square test is similar to the above. Finally, the dimensionality-reduced features are fused as the recognition basis for the subsequent compiler.
[0112] Step 5, Compiler recognition: Use four machine learning models, SVM, LightGBM, XGBoost, and RF, to conduct classification experiments on the dimensionality-reduced statistical features, associated features, and fused features on the 8 datasets in Step 4 to achieve the recognition of the compiler family, version, and optimization level.
[0113] The following conducts an analysis of the experimental results of four machine learning models and 7 sub-datasets for the compiler fusion feature extraction and optimization level recognition method of the present invention:
[0114] (1) Comparison of experimental results of four machine learning models
[0115] Use four machine learning models, SVM, LightGBM, XGBoost, and RF, to conduct classification experiments on the dimensionality-reduced single features and fused features respectively. The comparison of experimental accuracies is shown in Table 3.
[0116] Table 3 Comparison of experimental results of various features
[0117]
[0118] As can be seen from Table 3, the fusion features have the best recognition effect on 12 compilers, and LightGBM has the highest classification accuracy. Among them, when using statistical features for classification, the choice of machine learning model has a relatively large impact on the classification results. The classification accuracy of LightGBM for statistical features reaches 0.937; when using fusion features for classification, the choice of machine learning model has little impact on the classification results, and the classification accuracy of LightGBM is still the highest; on the premise of using the LightGBM machine learning model, the classification accuracy of fusion features is 3% higher than that of statistical features and 0.2% higher than that of correlation features. It can be seen that correlation features can also identify compilers well.
[0119] (2) Experimental results and analysis of 7 sub-datasets
[0120] The experimental results of the combination of LightGBM and fusion features are shown in Table 4. As can be seen from Table 4, the combination of chi-square, LightGBM and fusion features has good recognition effects on the family, version of the compiler and two optimization levels (binary classification of O0 and O2 optimization levels). The recognition accuracy for the four optimization levels reaches 0.984 for gcc and 0.940 for clang.
[0121] Table 4 Performance results on different datasets
[0122]
[0123]
[0124] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for compiler fusion feature extraction and recognition based on disassembly, characterized in that, The implementation is specifically carried out according to the following steps: Step 1, dataset construction; Step 2, data preprocessing: Use the IDA Pro disassembly tool to disassemble the executable file into a disassembly file, which is the asm file; Step 3, feature extraction: Extract statistical features and correlation features from the disassembly file obtained in Step 2. The statistical features include register usage frequency and opcode usage frequency; Step 4, feature preprocessing: Perform dataset partitioning on the statistical features and correlation features extracted in Step 3 to obtain 8 datasets. Use the chi-square test feature selection method to screen the extracted statistical features and correlation features respectively, and select the feature set with the top 40% chi-square scores. The screened statistical features and correlation features are fused as the recognition basis for the subsequent compiler; Step 5, compiler recognition: Use four machine learning models, SVM, LightGBM, XGBoost, and RF, to conduct classification experiments on the dimensionality-reduced statistical features, correlation features, and fused features on the 8 datasets in Step 4 to achieve the recognition of the compiler family, version, and optimization level; The specific implementation of Step 1 is carried out according to the following steps: Step 1.1, Download the CSmith test tool in the Linux operating system and configure the CSmith environment to generate 10,000 C language source codes as the dataset, denoted as the Dataset1 dataset; Step 1.2, In order to evaluate the extracted features from three perspectives: the compiler family, version, and optimization level, where the optimization levels are O0, O1, O2, and O3, download three types of compilers, gcc8.1.0, gcc9.2.0, and clang10.0.0, in the Win10 operating system; Step 1.3, Write a python batch script to compile the 10,000 C language source code files in the Dataset1 dataset in Step 1.1 into corresponding executable code files using the four optimization levels of the three compilers in Step 1.2, where the optimization levels are O0, O1, O2, and O3; 10,000 executable files are generated for each optimization level, and finally 120,000 executable files are obtained; Use the 120,000 executable files as the subsequent experimental dataset, denoted as the Dataset2 dataset, the categories and quantities of the Dataset2 dataset; The specific implementation of Step 2 is carried out according to the following steps: Preprocess the executable files in the Dataset2 dataset; Write an IDA Python automation script analysis program to call the IDA Pro tool to batch decompile the Dataset2 dataset in Step 1.3 into disassembly files with the suffix.asm, and denote this disassembly dataset as Data. Then extract statistical features and correlation features based on Data; The specific implementation of Step 3 is carried out according to the following steps: For the convenience of subsequent processing, store the features in a csv file after fusing them according to the labels; In the csv file, the first n columns are statistical features, and the remaining columns are correlation features; Step 3.1, Extract the register usage frequency from the disassembly file obtained in Step 2; Step 3.2: Extract the opcode usage frequency from the disassembly file obtained in Step 2. Step 3.3: Collectively refer to the register usage frequency extracted in Step 3.1 and the opcode usage frequency extracted in Step 3.2 as statistical features. Considering that the statistical features lack local information of the assembly code, an N-gram feature extraction method based on the statistical feature sequence is adopted to obtain correlation features, where N = 2, and N-gram refers to a continuous subsequence composed of N feature units. The specific implementation of Step 4 is as follows: Step 4.1: Feature set re-partitioning: The statistical features and correlation features extracted from the Dataset2 dataset are fused using the merge function in the panda library and stored in a CSV file. The statistical features and correlation features are divided into 7 sub-datasets according to the compiler's family, version, and optimization level, and finally 8 datasets are generated. The 8 datasets are: Dataset2 and 7 sub-datasets. Step 4.2: Since there are a large number of redundant features in the statistical features and correlation features extracted in Step 3.3, in order to improve the accuracy and training speed of model training, the chi-square test feature selection method is used to select the extracted statistical features and correlation features on the 8 datasets respectively. Select the feature set with the top 40% chi-square scores, and fuse the dimension-reduced features as the basis for compiler recognition later. The statistical features are described below. Step 4.3: Feature fusion: Fuse the statistical features and correlation features with the same label. Finally, a total of 972-dimensional features are extracted from each asm file and stored in a CSV file. The first 134 dimensions are statistical features, the first 42-dimensional features are register features, the last 92-dimensional features are opcode features, and the remaining 838 columns are correlation features. Then, first divide the extracted statistical features and correlation features into training sets and test sets, and then use the chi-square test feature selection method to select each feature according to the column where the feature is located, and fuse the selected features.
2. The method for extracting and identifying compiler fusion features based on disassembly according to claim 1, wherein The specific content of Step 3.1 is as follows: Use the usage frequencies of 42 registers as a feature of the compiler, and count the frequencies of register occurrences in each asm file. 42-dimensional register features are extracted from each asm file.
3. The method for compiler fusion feature extraction and recognition based on disassembly according to claim 1, wherein The specific content of Step 3.2 is as follows: Count the usage frequencies of 92 common opcodes in each asm file.
4. The method for extracting and identifying compiler fusion features based on disassembly according to claim 1, characterized in that The specific N-gram feature extraction method based on the statistical feature sequence in Step 3.3 is as follows: (1) Read the disassembly file in the Data dataset line by line, and extract the registers and opcodes that appear in the statistical feature sequence in the file. (2) Store the registers and opcodes obtained from the disassembly file in step (1) in a list, take every two elements in the list as a tuple, and count the frequencies of all pairs in the entire list in the Data dataset. Use the tuple as the key of the dictionary, and the frequency of the tuple occurrence as the value of the dictionary. The key composed of the pair is the N-gram feature based on the statistical features extracted from an disassembly file. (3) Screen the N-gram features of the entire dataset, select the features with a frequency greater than 2500, and finally extract 838-dimensional N-gram features from each disassembly file.
5. The method for extracting and identifying compiler fusion features based on disassembly according to claim 1, characterized in that The chi-square test feature selection method in step 4.2 is specifically as follows: Assume that the sample space is represented by the matrix X m×n =(x1, x2, …, x n , y), where m represents the total number of samples, n represents the number of features extracted from each sample, x i represents the i-th feature, and y represents the sample category; there are 8 data sets with different numbers of samples, so here m is used to represent the number of samples in the data set, n represents the feature category, and n is 134; (1) Assume that each feature is independent of the label y; (2) Calculate the chi-square value χ of each feature x i and label y using formula (3-1). 2 (x i , y); (3) Sort the n χ 2 (x1, y), χ 2 (x2, y), …, χ 2 (x n , y) in descending order and select features according to the set threshold; where χ 2 (x, y) represents the chi-square value calculated for each feature, x i represents the actual value of the i-th data; E i represents the expected value of the i-th data; n represents the total number of observed data; (4) The process of using the chi-square test to reduce the dimension of the associated features is similar to the above. Finally, the features after dimension reduction are fused as the recognition basis for the subsequent compiler.
Citation Information
Patent Citations
A malicious code classification method based on multiple features and feature selection
CN109033833A
Artificial intelligence mobile integration
US20200150933A1