Malware comparative learning detection method based on triple network

By using triple network structure and loss function in malware detection, the problem of insufficient malware detection performance in the prior art is solved, and stronger feature learning and category distinction capabilities are achieved, and robustness is enhanced.

CN120034383APending Publication Date: 2025-05-23BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187949.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

Smart Images

  • Figure CN120034383A_ABST
    Figure CN120034383A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security, and particularly relates to a malicious software comparative learning detection method based on a triple network. Comprising the following steps: (1) screening malicious software samples from a data set, carrying out feature extraction on the malicious software samples, and constructing a triple data set by utilizing the extracted features; (2) constructing a triple network, and performing feature embedding, loss function calculation and distance measurement; (3) inputting the triple data set into a triple network for training; and (4) inputting a new malicious software sample, extracting the feature embedding of the new malicious software sample by using the trained triple network, and simultaneously carrying out classification detection on the new malicious software sample. According to the method, by introducing the triple network structure, feature embedding of malicious software samples is learned more effectively, the detection capability of multi-class malicious software families is improved, the detection performance of minority-class malicious software families is enhanced, and the robustness of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a malware comparative learning detection method based on a triplet network. Background Art

[0002] With the popularization of mobile Internet, Android systems have become the main target of malware attacks. The diversity and variants of malware are increasing, causing traditional detection methods (such as feature matching) to gradually become ineffective. In recent years, malware detection methods based on deep learning have received widespread attention. Among them, the twin neural network has made some progress in dealing with data imbalance problems through contrastive learning, but there is still a defect of insufficient detection performance for a few types of malware families.

[0003] In the prior art, the contrastive learning method based on the twin neural network inputs a pair of malware samples (positive sample pair or negative sample pair) and trains the model to learn the similarities between samples. This method shows certain advantages in dealing with data imbalance problems, but in multi-category malware family detection, due to limited learning signals, it is difficult to fully capture the complex relationship between categories.

[0004] In summary, existing malware detection methods have the following technical problems:

[0005] 1. Limited learning signals: In the contrastive learning method based on the twin neural network, the twin neural network only considers the relationship between two samples at a time, which makes it difficult to fully learn the complex structure of multi-category malware families.

[0006] 2. Insufficient detection performance for minority malware families: In the contrastive learning method based on the twin neural network, in the scenario of data imbalance, the detection effect of the twin neural network on minority malware families is still not ideal.

[0007] 3. Insufficient robustness: In the contrastive learning method based on the twin neural network, the robustness of the twin neural network needs to be improved when facing malware variants and adversarial attacks. Summary of the invention

[0008] Therefore, in order to solve the technical problems of limited learning signals, insufficient detection performance for a few types of malware families, and insufficient robustness in the existing contrastive learning method based on twin neural networks, the present invention provides a malware contrastive learning detection method based on a triplet network.

[0009] The present invention aims to propose a malware comparative learning detection method based on triplet network. By introducing the triplet network structure, the feature embedding of malware samples can be learned more effectively, the detection capability of multi-category malware families can be improved, the detection performance of minority malware families can be enhanced, and the robustness of the model can be improved.

[0010] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0011] The present invention provides a malware contrast learning detection method based on a triplet network, which mainly includes the following steps:

[0012] (1) Filter malware samples from the dataset, extract features from them, and construct a triplet dataset using the extracted features;

[0013] (2) Construct a triplet network and perform feature embedding, loss function calculation, and distance measurement;

[0014] (3) Input the triplet data set into the triplet network for training;

[0015] (4) Input a new malware sample, use the trained triplet network to extract its feature embedding, and perform classification detection on the new malware sample.

[0016] Furthermore, the dataset is the AndroZoo dataset.

[0017] Furthermore, in step (1), the static analysis tool Androguard is used to extract the static features of the API call sequence and the static features of the permission request of the malware sample; the dynamic analysis tool DroidBox is used to extract the dynamic features of the network connection and the dynamic features of the file operation of the malware sample.

[0018] Furthermore, the triplet data set includes: anchor samples, positive samples and negative samples; the anchor samples are randomly selected malware samples; the positive samples are samples belonging to the same malware family as the anchor samples; the negative samples are samples belonging to different malware families than the anchor samples.

[0019] Furthermore, the triplet network is composed of three sub-networks with shared weights, each of which uses a multi-layer perceptron; the multi-layer perceptron includes multiple hidden layers, and the hidden layers use a ReLU activation function to extract features of the samples and embed the features into a continuous vector space.

[0020] Furthermore, the loss function is a triplet loss function, and its calculation formula is:

[0021] L=max(0,d anchor-positive -danchor-negative +α) (1)

[0022] Among them, d anchor-positive and d anchor-negative They represent the Euclidean distance between the anchor sample and the positive sample, and between the anchor sample and the negative sample respectively. α is a hyperparameter used to control the distance interval.

[0023] Furthermore, the Euclidean distance is used as the distance metric between samples, and its calculation formula is:

[0024]

[0025] Among them, x and y represent the feature vectors of positive samples and negative samples respectively, D represents the dimension of the feature vector, and x i Represents the eigenvalue of the positive sample in the i-th dimension, y i Represents the eigenvalue of the negative sample in the i-th dimension.

[0026] Furthermore, during the triplet network training process, the network parameters are updated through back-propagation to minimize the loss function and learn the feature embedding of malware samples.

[0027] Furthermore, during the training of the triplet network, a hard negative sample mining strategy is adopted to select negative samples that are closest to the anchor samples to enhance the model's discrimination ability.

[0028] Furthermore, the classification detection method is:

[0029] First, we calculate the distance between the new malware sample and the center of the known malware family, and select the malware family closest to it as the prediction result to determine the malware family it belongs to.

[0030] Then, by setting a threshold, we determine whether the new malware sample belongs to an unknown malware family. If the distance between the new malware sample and all known malware families is greater than the threshold, it will be marked as an unknown malware family to achieve malware classification detection.

[0031] The beneficial effects of the present invention are:

[0032] 1. Stronger feature learning ability: Through comparative learning of triplet networks, the present invention can more effectively learn the feature embedding of malware samples and improve detection performance.

[0033] 2. Better category differentiation capability: The present invention can better handle the detection problem of multiple categories of malware families, especially significantly improving the detection performance of a few categories of malware families.

[0034] 3. Higher robustness: The present invention has higher robustness against malware variants and adversarial attacks, and can more stably cope with complex security threats. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 The input and output of the triplet network and the calculation process of the loss function. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0037] The present invention provides a malware contrast learning detection method based on a triplet network (TripletNetwork), which is mainly used for Android malware detection. In Android malware detection, the data imbalance problem seriously affects the performance and generalization ability of the detection model, especially the poor detection effect for a few malware families. Although the existing technology (such as the contrast learning method based on the twin neural network) has alleviated the data imbalance problem to a certain extent, it still has limitations when dealing with multi-category malware families, and it is difficult to effectively distinguish a few classes of malware families. In addition, the existing technology is not robust enough to malware variants and adversarial attacks, resulting in a decrease in detection performance. Therefore, a more efficient technical solution is needed to improve the accuracy and robustness of malware detection, while better coping with the diversity and complexity of malware families. The present invention trains the network to learn the feature embedding of malware samples by inputting an anchor sample (Anchor), a positive sample (Positive) and a negative sample (Negative), so that samples of the same malware family are closer in the feature space and samples of different malware families are farther away; the network parameters are optimized by the triplet loss function, the detection ability of the model for multi-category malware families is improved, the detection performance for a few types of malware families is enhanced, and the robustness of the model is improved.

[0038] The present invention provides a malware comparative learning detection method based on a triplet network, and its specific implementation process is as follows:

[0039] Step S1: data preparation;

[0040] S1.1: Data sources;

[0041] We use the AndroZoo dataset, which contains a large number of malware samples covering Android apps from 2010 to the present.

[0042] In the present invention, malware samples from the AndroZoo dataset over the past five years are screened to ensure that the features learned by the model are similar to modern malware.

[0043] S1.2: Data preprocessing;

[0044] Perform feature extraction on malware samples, including static features (such as API call sequences, permission requests) and dynamic features (such as network behavior, file operations).

[0045] Specifically, static analysis tools (such as Androguard) can be used to extract static features such as API call sequences and permission requests of malware samples.

[0046] Specifically, dynamic analysis tools (such as DroidBox) can be used to extract dynamic features such as network connections and file operations of malware samples.

[0047] S1.3: Triple dataset construction;

[0048] The extracted features are used to construct triple datasets, each of which contains:

[0049] Anchor sample: a randomly selected malware sample;

[0050] Positive: A sample that belongs to the same malware family as the anchor sample.

[0051] Negative samples: samples that belong to different malware families than the anchor samples.

[0052] The present invention needs to ensure that the selection of positive and negative samples is representative to cover the diversity and complexity of malware families.

[0053] At the same time, the present invention uses data enhancement technology (such as random noise injection) to expand minority malware samples to improve the sensitivity of the model to minority malware samples.

[0054] Step S2: triplet network model;

[0055] Compared with the twin network, the triplet network can more effectively learn the feature embedding of malware samples and improve the detection performance by simultaneously considering the relationship between anchor samples, positive samples and negative samples.

[0056] Figure 1 The input (anchor samples, positive samples, negative samples) and output (feature embedding) of the triplet network and the calculation process of the loss function are shown in Figure 2.

[0057] S2.1: Triplet network architecture;

[0058] The present invention adopts a triplet network model, which consists of three sub-networks with shared weights, each of which is responsible for processing an input sample (anchor sample, positive sample, negative sample). Each sub-network adopts a multi-layer perceptron (MLP) structure, including multiple hidden layers, to extract the features of the sample and embed the features into a continuous vector space.

[0059] S2.2: Feature embedding;

[0060] After the input sample passes through the MLP network, the output is a feature vector of fixed dimension, which represents the embedding of the sample in the feature space.

[0061] The dimension of feature embedding is selected according to experimental results, such as 128 or 256 dimensions.

[0062] In addition, each hidden layer in the present invention adopts the ReLU activation function to introduce nonlinear characteristics and enhance the expressive power of the model.

[0063] S2.3: Loss function calculation;

[0064] The triplet loss function is used to optimize the network parameters to ensure that the distance between the anchor sample and the positive sample is smaller than the distance between the anchor sample and the negative sample.

[0065] The specific calculation formula of the triplet loss function is:

[0066] L=max(0,d anchor-positive -d anchor-negative +α) (1)

[0067] Among them, d anchor-positive and d anchor-negative They represent the Euclidean distance between the anchor sample and the positive sample, and between the anchor sample and the negative sample respectively. α is a hyperparameter used to control the distance interval.

[0068] The present invention ensures the clustering of the same malware family and the distinguishability of different malware families in the feature space by using a triplet loss function.

[0069] S2.4: distance metric;

[0070] The Euclidean distance is used as the distance metric between samples. The specific calculation formula is:

[0071]

[0072] Among them, x and y represent the feature vectors of positive samples and negative samples respectively, D represents the dimension of the feature vector, and x iRepresents the eigenvalue of the positive sample in the i-th dimension, y i Represents the eigenvalue of the negative sample in the i-th dimension.

[0073] Step S3: training;

[0074] (1) Training process

[0075] Input samples: Input anchor samples, positive samples and negative samples into the triplet network respectively;

[0076] Feature extraction: Each sub-network extracts feature embedding of the input sample;

[0077] Loss calculation: Calculate the distance between the anchor sample and the positive sample, the distance between the anchor sample and the negative sample, and calculate the loss value through the triplet loss function.

[0078] Parameter Update: During the training process, the network parameters are updated through back-propagation to minimize the triplet loss function and learn the feature embedding of the malware samples.

[0079] The present invention introduces adversarial samples during the training process, further enhancing the robustness of the model.

[0080] (2) Optimization strategy

[0081] The Adam optimizer is used to update parameters, and the learning rate is adjusted according to the experimental results.

[0082] During the training process, the performance of the model on the validation set is regularly evaluated, and hyperparameters (such as learning rate, hyperparameter α) are adjusted to optimize model performance.

[0083] Among them, the early stopping mechanism can be used to prevent overfitting. When the loss value on the validation set no longer decreases within multiple consecutive epochs, the training is stopped.

[0084] (3) Training details

[0085] Batch Size: Select an appropriate batch size based on your hardware resources, such as 32 or 64.

[0086] Learning rate adjustment: Adopt the learning rate decay strategy to gradually reduce the learning rate as the training progresses to improve the convergence speed and accuracy of the model.

[0087] Positive and negative sample selection strategy: The hard negative mining strategy is used to select negative samples that are closest to the anchor samples to enhance the model's ability to distinguish.

[0088] Through comparative learning, the model can better capture the characteristic differences of malware families and improve detection performance. At the same time, the model is more robust to malware variants and adversarial attacks and can more stably respond to complex security threats. In addition, the model can better handle the detection problems of a few types of malware families and improve overall detection performance.

[0089] Step S4: detection;

[0090] S4.1: Feature extraction;

[0091] For new malware samples, the trained triplet network is used to extract their feature embedding; after the input sample passes through the MLP network, the output is a feature vector of fixed dimension.

[0092] S4.2: Classification test;

[0093] First, by calculating the distance between the new malware sample and the center of the known malware family, the malware family closest to it is selected as the prediction result to determine the malware family it belongs to.

[0094] Then, by setting a threshold, we determine whether the new malware sample belongs to an unknown malware family. If the distance between the new malware sample and all known malware families is greater than the threshold, it will be marked as an unknown malware family, ultimately achieving malware classification detection.

[0095] S4.3: Real-time optimization;

[0096] In order to improve the real-time performance of the detection system, the trained model is quantified and optimized to reduce the computational complexity of the model.

[0097] Among them, a lightweight network structure (such as MobileNet) can be used to replace part of the MLP network to reduce computing resource consumption.

[0098] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A malware contrast learning detection method based on triplet network, characterized in that: The following steps are involved: (1) Filter malware samples from the dataset, extract features from them, and construct a triplet dataset using the extracted features; (2) Construct a triplet network and perform feature embedding, loss function calculation, and distance measurement; (3) Input the triplet data set into the triplet network for training; (4) Input a new malware sample, use the trained triplet network to extract its feature embedding, and perform classification detection on the new malware sample.

2. According to the triplet network-based malware contrast learning detection method of claim 1, it is characterized in that: The dataset is the AndroZoo dataset.

3. The malware contrast learning detection method based on triplet network according to claim 1 is characterized in that: In step (1), the static analysis tool Androguard is used to extract the static features of the API call sequence and the static features of the permission request of the malware sample; The dynamic analysis tool DroidBox is used to extract the network connection dynamic features and file operation dynamic features of the malware samples.

4. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: The triplet data set includes: anchor samples, positive samples and negative samples; the anchor samples are randomly selected malware samples; the positive samples are samples belonging to the same malware family as the anchor samples; the negative samples are samples belonging to different malware families than the anchor samples.

5. The malware contrast learning detection method based on triplet network according to claim 1 is characterized in that: The triplet network is composed of three sub-networks with shared weights, each of which uses a multi-layer perceptron; the multi-layer perceptron includes multiple hidden layers, and the hidden layers use a ReLU activation function to extract the features of the samples and embed the features into a continuous vector space.

6. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: The loss function is a triplet loss function, and its calculation formula is: L=max(0,d anchor-positive -d anchor-negative +α) (1) Among them, d anchor-positive and d anchor-negative They represent the Euclidean distance between the anchor sample and the positive sample, and between the anchor sample and the negative sample respectively. α is a hyperparameter used to control the distance interval.

7. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: The Euclidean distance is used as the distance metric between samples, and its calculation formula is: Among them, x and y represent the feature vectors of positive samples and negative samples respectively, D represents the dimension of the feature vector, and x i Represents the eigenvalue of the positive sample in the i-th dimension, y i Represents the eigenvalue of the negative sample in the i-th dimension.

8. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: During the triplet network training process, the network parameters are updated through back-propagation to minimize the loss function and learn the feature embedding of malware samples.

9. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: During the training process of the triplet network, a hard negative sample mining strategy is adopted to select negative samples that are closest to the anchor samples to enhance the model's discrimination ability.

10. The malware contrast learning detection method based on triplet network according to claim 1, characterized in that: The classification detection method is: First, we calculate the distance between the new malware sample and the center of the known malware family, and select the malware family closest to it as the prediction result to determine the malware family it belongs to. Then, by setting a threshold, we determine whether the new malware sample belongs to an unknown malware family. If the distance between the new malware sample and all known malware families is greater than the threshold, it will be marked as an unknown malware family to achieve malware classification detection.