A Malicious Sample Detection Method under the Condition of Unbalanced Class Sample Quantities

Through feature extraction and similarity calculation, the optimal sampling parameter group was selected, combined with the LightGBM model, the problem of category imbalance in malicious sample detection was solved, and the downsampling of most classes and oversampling of few classes was realized, which improved the detection generalization ability.

CN114548305BActive Publication Date: 2025-07-29BEIJING VENUS INFORMATION SECURITY TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210187808.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-29
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In malicious sample detection, it is difficult for the prior art to effectively utilize unlabeled data, and traditional methods cannot optimize sampling parameters under category imbalance, resulting in poor classifier effects.

Method used

Through feature extraction, classification algorithm prediction result similarity calculation and selection of optimal sampling parameter groups, combined with the LightGBM model, downsampling of most classes and oversampling of few classes are realized to optimize training data.

Benefits of technology

The generalization ability of malicious sample detection is improved, and the sampling parameters are optimized using labelless data to achieve balanced sampling of most and few classes, which improves the overall effect of the classifier.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548305B_ABST
    Figure CN114548305B_ABST
Patent Text Reader

Abstract

The present application provides a malicious sample detection method in the case of unbalanced class sample quantities. The steps include: extracting features from the original samples with unbalanced class sample quantities to obtain the samples after feature extraction as training data; using a classification algorithm to obtain at least two classification prediction results of the training data; wherein, the training data includes unlabeled data; setting a set of sampling parameter groups, the set of sampling parameter groups being composed of several sampling parameter groups, and each sampling parameter group including sampling parameters used when sampling samples of various classes in the training data; taking the sampling parameter group in the set of sampling parameter groups that makes the similarity between all classification prediction results the highest as the optimal sampling parameter group; and sampling the training data according to the optimal sampling parameter group. Using the present application can simultaneously downsample the majority class and oversample the minority class, thereby improving the generalization ability of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network information security technology, and particularly relates to a malicious sample detection method for solving classification problems under the condition of class imbalance. Background Art

[0002] In recent years, malicious files of different families, such as ransomware, Trojans, viruses, mining programs and other malicious software, have emerged continuously, bringing a lot of trouble and economic losses to users and institutions. In order to improve the efficiency of analyzing a large number of malicious software, it is necessary to distinguish the families of malicious software. In the application of network security classifiers, the vast majority of data scenarios encountered are those with imbalanced sample classes. Among them, we set the class with a large number of samples as the majority class, and the class with a small number of samples as the minority class. In the research, it can be found that under normal circumstances, the amount of data in the majority class will far exceed that in the minority class. Under such a huge data volume gap, if the data is not adjusted, it will cause the classifier to be biased towards the majority class, resulting in a deterioration of the overall effect of the classifier. In the scenario of data imbalance, it can also be found that although the majority class has a large number of samples, it often contains a large number of redundant samples, while the minority class, although having a small number of samples, often contains more information.

[0003] With the continuous development of science and technology and the continuous change of the security situation, the number of malicious samples has increased exponentially, and there are endless variants. Therefore, the detection of malicious samples will play an increasingly important role. The traditional detection method based on artificial rules has a relatively low development efficiency and poor generalization. In order to better summarize and learn a large amount of data, machine learning methods are generally used in the existing technology, which can improve the generalization ability of detection.

[0004] In the scenario of malicious sample detection based on machine learning, classifiers are mainly divided into binary classification of normal samples and malicious samples and multi-classification of different malicious sample families. Among them, in binary classification, the problem that is likely to occur is that the proportion of normal samples is relatively large, while the proportion of malicious samples is relatively small. In multi-classification, the problem that is likely to occur is that the proportion of samples in some families is relatively large, while the samples in some individual families are very few. Thus, it can be seen that the problem of class imbalance often exists in malicious sample identification. One of the mainstream methods to solve class imbalance in classification problems is data sampling. Specifically, it is to oversample the minority class samples and downsample the majority class samples. However, the existing methods can only sample on the training set and then select by comparing with other labeled data. In actual application scenarios, there are often a large number of unlabeled data, and the above methods cannot make good use of the unlabeled data. At the same time, due to the inability to provide hyperparameters for sampling that can balance and optimize the number of class samples after sampling, the classification effect is poor. Summary of the Invention

[0005] To solve the above problems, the present application provides a malicious sample detection method in the case of unbalanced class sample numbers, and the steps include:

[0006] S1. Extract features from the original samples with unbalanced class sample numbers to obtain the samples after feature extraction as training data;

[0007] S2. Use a classification algorithm to obtain at least two classification prediction results of the training data; wherein, the training data includes unlabeled data;

[0008] Set a set of sampling parameter groups, which is composed of several sampling parameter groups, and each sampling parameter group includes sampling parameters used for sampling various class samples in the training data;

[0009] Among the set of sampling parameter groups, take the sampling parameter group with the highest similarity between all classification prediction results as the optimal sampling parameter group;

[0010] S3. Sample the training data according to the optimal sampling parameter group and train the sampled samples.

[0011] Among them, preferably, in step S2, the classification algorithm includes the K-nearest neighbor algorithm, and the K-nearest neighbor algorithm can determine the class of a sample according to the class to which the majority of the K nearest instances belong.

[0012] Among them, preferably, in step S2, it further includes:

[0013] S21. Obtain the structural similarity Q between classification prediction results m ;

[0014] S22. Obtain the distribution similarity Q between classification prediction results n ;

[0015] S23. Obtain the similarity Q between classification prediction results according to the structural similarity Q m and the distribution similarity Q m ;

[0016] Among them, preferably, the calculation method of the structural similarity Q m is:

[0017] Sort the classes of the classification prediction results in descending order according to the number of samples included as the classification ranking of the classification prediction results; according to the classification ranking, set the weights of the corresponding classes from high to low;

[0018] Compare the classification rankings of all classification prediction results. When the same class appears at the same sequence position of different classification prediction results, set the class as a similar class;

[0019] Set the sum of the weights of all categories in the classification prediction result to ∑W j , and the sum of the weights of similar categories to ∑W i ;

[0020] Then the structural similarity can be obtained

[0021] Among them, preferably, the distribution similarity Q n is calculated as follows:

[0022] Set the sample set in the r-th category of the first classification prediction result as R1, the sample set in the r-th category of the second classification prediction result as R2, and the set of overlapping sample IDs between R1 and R2 as R0.

[0023] Set the Jacaard similarity coefficient of the r-th category samples of the first prediction result and the second prediction result

[0024]

[0025] Set there are δ categories, then the distribution similarity of the first prediction result and the second prediction result is obtained

[0026]

[0027] Among them, preferably, in step S23, the similarity Q between all classification prediction results is the structural similarity Q m and the distribution similarity Q n is the weighted sum of.

[0028] Among them, preferably, in step S1, it further includes:

[0029] Step S11, extract visible characters from the original sample and disassemble the original sample file to obtain a disassembly file;

[0030] Step S12, automatically screen out important strings through a search engine, extract the N-gram features of the common instructions in the disassembly file, and merge the two as the final features.

[0031] Among them, preferably, in step S3, after sampling, the LightGBM model is used for training, and finally a trained model is obtained.

[0032] Among them, preferably, in step S2, according to the imbalance of the number of category samples, a sampling parameter set is set. Among them, the sampling parameter for the majority class samples is set to be greater than 0 and less than 1.0, while the sampling parameter for the minority class samples is set to be greater than 1 and less than the number of samples of the most numerous class divided by the number of categories.

[0033] The beneficial effects achieved by this application are as follows:

[0034] Using the method of the present application can make good use of unlabeled data, conveniently and quickly obtain relatively optimized sampling parameters, can simultaneously downsample the majority class and oversample the minority class, so that the majority class and the minority class achieve a balanced sampling effect, thereby improving the generalization ability of detection. Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those skilled in the art, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a flowchart for implementing the malicious sample detection method in the case of unbalanced class sample numbers of the present application in modules. Detailed Embodiments

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0038] The present application provides a malicious sample detection system in the case of unbalanced class sample numbers, including a feature extraction module, a sampling parameter setting module, a sampling parameter search module, and a similarity calculation module.

[0039] Among them, the feature extraction module extracts features from the original sample file. First, visible characters are extracted from the original sample and the original sample file is disassembled to obtain an assembly file. Then, important strings are automatically screened out through a search engine, and the N-gram features of the common instructions in the assembly file are extracted. The two features are combined as the final feature.

[0040] In the sampling parameter setting module, it is possible to select to automatically generate m groups of sampling parameters or manually fill in n groups of sampling parameters. Among the automatically generated m groups of sampling parameters, the sampling ratio for the majority class samples is an array greater than 0 and less than 1.0, such as [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9], while the sampling ratio for the minority class is an array greater than 1 and less than the maximum number of majority class samples divided by the number of this class.

[0041] The sampling parameter search module is used to obtain all classification prediction results of the training data using a classification algorithm; a set of sampling parameter groups is set, and the set of sampling parameter groups is composed of several sampling parameter groups, and the sampling parameter group includes sampling parameters used when sampling samples of each category in the classification prediction results; when the p-th parameter group in the sample sampling parameter group set is used, the similarity between all classification prediction results of the training data is the highest, then the p-th parameter group is used as the optimal sampling parameter of the training data.

[0042] For example, use K-Nearest Neighbor (KNN) with different parameters to fit the training set, and then predict the unlabeled data to obtain the prediction results of each unlabeled data. Calculate the similarity between pairwise different prediction results, and divide the sum by the number of combinations between pairwise to obtain the score corresponding to each set of parameters. Use the set of parameters with the highest score as the optimal parameter after the search.

[0043] Table 1 shows two classification prediction results A and B of the training data in a specific embodiment, where the sample categories are represented by 0 to 8.

[0044]

[0045]

[0046] Table 1

[0047] In the similarity calculation module, the similarity refers to the weighted sum of the structural similarity and the distribution similarity. Sort the categories of the classification prediction results in descending order according to the number of samples included, as the classification sorting of the classification prediction results; according to the classification sorting, set the weights corresponding to the categories from high to low; compare the classification sortings of all classification prediction results, and when the same category appears at the same sequence position of different classification prediction results, set the category as the similar category. For example, the weights corresponding to the two classification results in Table 1 after sorting are 1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 respectively, then the structural similarity between the A prediction result and the B prediction result is = (1 + 0.9 + 0.5 + 0.4 + 0.3 + 0.2) / (1 + 0.9 + 0.8 + 0.7 + 0.6 + 0.5 + 0.4 + 0.3 + 0.2 + 0.1) = 0.6. Note: If it matches, add the weight, otherwise add 0.

[0048] The distribution similarity refers to calculating the Jacaard coefficient (the number of intersections divided by the number of unions) corresponding to each category position respectively, and dividing the sum by the number of categories. For example, set the sample set in the r-th category of the first classification prediction result as R1, the sample set in the r-th category of the second classification prediction result as R2, and the set of sample IDs that overlap between R1 and R1 as R0.

[0049] Set the Jacaard similarity coefficient between the first prediction result and the second prediction result as J i , then the Jacaard similarity coefficient of the r-th category is obtained

[0050] Set that there are δ categories in total, then the distribution similarity between the first prediction result and the second prediction result is obtained

[0051]

[0052] The model training module samples the training data using the set of parameters with the highest score, and after sampling, uses the LightGBM model for training, and finally obtains the trained model

[0053] For example, in a multi-classification problem of malicious samples, it is necessary to classify the families of malware, where the categories are identified from 0 to 8. Among them, 50,000 are the training set, and 8,000 are used as the test set. The evaluation metric for the test set is macro F1-score. The category distribution of the training set is shown in Table 2

[0054] Category 7 8 5 6 3 4 2 1 0 Number of samples 11368 9425 7563 7560 6641 5676 784 598 385

[0055] Table 2

[0056] According to the category distribution, it can be seen that the most numerous category is the 7th category, and the least numerous category is the 0th category, and the ratio between the two is about 30 times, so there is an obvious imbalance between the categories. After preprocessing and feature extraction of the training data, then sampling is used to deal with the imbalance phenomenon, and finally the LightGBM model is used for classification

[0057] The different classification results obtained according to different sampling ratios are shown in Table 3, where F1 is the accuracy judgment of the classification results by a third-party evaluation agency

[0058]

[0059] Table 3

[0060] According to Table 3, it can be seen that the sampling parameter group with the highest similarity is [7th category: 500; 3rd category: 3000; 0th category: 500], and the accuracy of the classification results obtained by sampling according to this sampling parameter is also the highest

[0061] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application. Obviously, those skilled in the art can make various changes and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A malicious sample detection method in the case of unbalanced class sample quantities, characterized by the steps Including: S1. Extract features from the original samples with imbalanced class sample numbers to obtain the samples after feature extraction as training data; S2. Use a classification algorithm to obtain at least two classification prediction results of the training data; wherein, the training data includes unlabeled data; Set a set of sampling parameter groups, which is composed of several sampling parameter groups, and each sampling parameter group includes sampling parameters used for sampling samples of each category in the training data; Take the sampling parameter group in the set of sampling parameter groups that makes the similarity between all classification prediction results the highest as the optimal sampling parameter group; S3. Sample the training data according to the optimal sampling parameter group, and train the samples obtained by sampling.

2. The malicious sample detection method in the case of unbalanced class sample quantities as described in claim 1, wherein, In step S2, the classification algorithm includes the K-nearest neighbor algorithm, and the K-nearest neighbor algorithm can determine the category of a sample according to the category to which the majority of the K nearest instances belong.

3. The malicious sample detection method in the case of unbalanced class sample quantities as described in claim 1, characterized in that, In step S2, it further includes: S21, obtain the structural similarity Q between the classification prediction results m ; S22, obtaining the distribution similarity Q between the classification prediction results n ; S23, according to the structural similarity Q m and the distribution similarity Q m obtain the similarity Q between the classification prediction results.

4. The malicious sample detection method in the case of unbalanced class sample numbers as described in claim 3, characterized in that The structural similarity Q m is calculated as follows: Sort the categories of the classification prediction results in descending order according to the number of samples included as the classification ranking of the classification prediction results; according to the classification ranking, set the weights of the corresponding categories from high to low; Compare the classification rankings of all classification prediction results. When the same category appears at the same sequence position in different classification prediction results, set the category as a similar category; Set the sum of the weights of all categories in the classification prediction result to ∑W j , and the sum of the weights of similar categories to ∑W i ; The structural similarity can then be obtained 5. The malicious sample detection method in the case of unbalanced class sample numbers as described in claim 3, characterized in that, The distribution similarity Q n is calculated as follows: Set the sample set in the r-th category of the first classification prediction result as R1, the sample set in the r-th category of the second classification prediction result as R2, and the set of overlapping sample IDs between R1 and R2 as R0, Set the Jacaard similarity coefficient of the r-th category samples of the first prediction result and the second prediction result Suppose there are a total of δ categories, then the distribution similarity between the first prediction result and the second prediction result is obtained 6. The malicious sample detection method in the case of unbalanced class sample quantities as described in claim 3, characterized in that, In step S23, the similarity Q among all classification prediction results is the structural similarity Q m and the distribution similarity Q n is the weighted sum of them.

7. The malicious sample detection method in the case of unbalanced class sample quantities as described in claim 1, characterized in that, In step S1, it further includes: Step S11. Extract visible characters from the original samples and disassemble the original sample files to obtain disassembly files; Step S12. Extract the N-gram features of the common instructions in the disassembly files for the important strings automatically screened by the search engine, and merge the two as the final features.

8. The malicious sample detection method under the condition of unbalanced class sample quantity as described in claim 1, characterized in that In step S3, after sampling, use the LightGBM model for training to finally obtain a trained model.

9. The malicious sample detection method in the case of unbalanced class sample quantities as described in claim 1, characterized in that In step S2, set the set of sampling parameter groups according to the imbalance of class sample numbers. Among them, the sampling parameters for majority class samples are set to be greater than 0 and less than 1.0, while the sampling parameters for minority class samples are set to be greater than 1 and less than the number of samples of the majority class divided by the number of classes.

Citation Information

Patent Citations

  • Imbalanced data optimal learning sample compound algorithm selection and parameter determining method

    CN110021426A

  • Malicious traffic detection method in data imbalance scene

    CN112990286A