A method and system for updating an Internet of Things malicious code detection model

By deploying a lightweight graph convolutional model on the edge and combining knowledge distillation and incremental learning techniques, the latency, accuracy, and adaptability issues in IoT malicious code detection are solved, achieving efficient and real-time malicious code detection and model updates.

CN118827160BActive Publication Date: 2025-09-09GUANGXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410807145.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-09-09
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

Existing IoT malicious code detection methods have problems with latency and real-time performance, low detection accuracy, and poor adaptability to dynamic environments in the IoT environment. In particular, cloud computing-based methods lead to delays and insufficient detection accuracy of lightweight models, making it difficult for edge devices to be updated in a timely manner.

Method used

A lightweight graph convolutional model is deployed on the edge for preliminary detection, and the detection capability of the high-performance cloud model is transferred to the lightweight model through knowledge distillation technology. At the same time, incremental learning technology is used to update the cloud and edge models to achieve rapid response and efficient update of the model.

Benefits of technology

It solves the latency and real-time issues in the IoT environment, improves detection accuracy, and enhances adaptability to dynamic environments, ensuring that edge devices can update detection models in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827160B_ABST
    Figure CN118827160B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for updating a malicious code detection model of the Internet of Things, comprising: the edge end reads multiple source files from the local, and converts all the source files into multiple RGB images in batches, all the RGB images constitute an image set, the edge end inputs each RGB image in the obtained image set into a pre-trained first detection model to obtain a detection result of the source file corresponding to the RGB image, and judges whether the source file contains malicious code based on the detection result. If malicious code is contained, the source file is uploaded to the malicious code data center in the cloud for data update. The cloud obtains all updated malicious code samples, and uses the samples to update the pre-trained second detection model. The present invention can solve the problems of the existing method of using cloud computing to upload the file to be detected to the cloud and wait for feedback, resulting in delays and not being real-time enough, the edge detection model cannot be updated in time, and the adaptability to dynamic environments is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security technology, and more specifically, relates to a method and system for updating a malicious code detection model for the Internet of Things. Background Art

[0002] With the rapid development and widespread adoption of the Internet of Things (IoT), security issues surrounding IoT devices are becoming increasingly prominent. Due to their typically weak security protections and computing capabilities, IoT devices have become a prime target for cyber attackers. In this context, detecting and preventing malicious code is crucial to maintaining the security of IoT devices and networks.

[0003] Among existing malicious code detection methods, signature-based detection methods and behavior-based detection methods are the two most widely used methods. The former relies on a signature library of pre-defined malware features (such as file fingerprints, string patterns, binary code sequences, etc.) and detects by matching known malicious code features, but cannot detect new or variant malicious codes; the latter identifies potential malicious activities by monitoring program behavior on the device (such as file operations, network requests, etc.), but may generate more false alarms and consume a lot of system resources. Therefore, these two methods are difficult to implement in an IoT environment where attacks are highly variable and resources are limited.

[0004] In order to overcome the defects of the above two methods, IoT malicious code detection based on deep learning methods has been increasingly widely used. It includes two implementation methods. The first is to use cloud computing to deploy the deep learning detection model in the cloud. When the edge IoT device performs malicious code detection, it uploads the file to be detected to the cloud and waits for the return result; the second is based on a lightweight detection model. According to the performance of the edge IoT device, a lightweight detection model is deployed for real-time detection.

[0005] However, both of the above-mentioned IoT malicious code detection methods based on deep learning methods have some non-negligible flaws:

[0006] First, using cloud computing requires uploading the files to be tested to the cloud and waiting for feedback, which will cause delays and real-time issues;

[0007] Second, the lightweight detection model-based approach has relatively low detection accuracy due to its shallow network depth, making it difficult to learn the deep features of malicious code.

[0008] Third, the lightweight detection model-based approach is difficult to implement timely updates of detection models on edge IoT devices when the IoT environment is changing rapidly, devices are constantly updated, and malicious code is constantly evolving. This results in poor adaptability of this approach to dynamic environments. Summary of the Invention

[0009] In response to the above defects or improvement needs of the existing technology, the present invention provides a method and system for updating the Internet of Things malicious code detection model, which aims to solve the technical problems that the existing method of using cloud computing requires uploading the file to be detected to the cloud and waiting for feedback, which will cause delays and real-time problems, and the technical problems that the existing method based on lightweight detection models has relatively low detection accuracy due to the shallow network depth of the lightweight model and it is difficult to learn the deep-level features of the malicious code; and the technical problems that the existing method based on lightweight detection models has poor adaptability to dynamic environments due to the rapid changes in the Internet of Things environment, continuous updates of devices, and continuous evolution of malicious codes, which makes it difficult for edge Internet of Things devices to achieve timely updates of detection models.

[0010] To achieve the above objectives, according to one aspect of the present invention, a method for updating an IoT malicious code detection model is provided. The method is applied in an environment including a first detection model, a cloud-based malicious code data center, and a second detection model. The method includes the following steps:

[0011] (1) The edge reads multiple source files from the local machine and converts all the source files into multiple RGB images in batches. All the RGB images constitute an image set.

[0012] (2) The edge terminal inputs each RGB image in the image set obtained in step (1) into a pre-trained first detection model to obtain a detection result of the source file corresponding to the RGB image, and determines whether the source file contains malicious code based on the detection result. If malicious code is contained, the source file is uploaded to the malicious code data center in the cloud for data update and the process proceeds to step (3). Otherwise, the process ends.

[0013] (3) obtaining, in the cloud, the malicious code samples updated by the cloud malicious code data center after the data is updated in step (2), and using the malicious code samples to update the pre-trained second detection model, thereby obtaining an updated second detection model;

[0014] (4) The cloud updates the first detection model at the edge using the second detection model updated in step (3) to obtain an updated first detection model;

[0015] (5) The cloud determines whether it has received a termination instruction from the client. If so, the process ends; otherwise, it returns to step (1).

[0016] Preferably, step (1) specifically comprises: obtaining each source file and reading each byte data in the source file; then, mapping each three bytes of data to an RGB pixel point, wherein the first byte data represents the red component, the second byte data represents the green component, and the third byte data represents the blue component; then, filling the obtained multiple RGB pixel points into the image to obtain an intermediate image, wherein the height and width of the intermediate image are determined according to the size of the source file; finally, using bilinear interpolation to scale the size of the obtained intermediate image to 128*128 to obtain the final RGB image.

[0017] Preferably, the first detection model is a lightweight graph convolution model, which specifically includes an initial part, four basic modules, three subsampling layers and a final part. The specific structure of the model is:

[0018] The initial part takes an image of dimension 3*128*128 as input and first processes it using a 3*3 regular convolution with a stride of 2 to obtain a first feature tensor of dimension 16*64*64. Next, a 3*3 depthwise separable convolution (DW-Conv) is used to process the feature tensor to obtain a second feature tensor of dimension 32*64*64. Subsequently, a 1*1 pointwise convolution operation is performed on the obtained second feature tensor to obtain a third feature tensor. Finally, the third feature tensor is further sampled using a 3*3 depthwise separable convolution with a stride of 2 to obtain a fourth feature tensor of dimension 32*32*32. In addition, the initial part also adds a shortcut connection between the output of the first layer 3*3 regular convolution and the output of the 1*1 pointwise convolution.

[0019] The first basic module, whose input is the fourth feature tensor of dimension 32*32*32 output by the starting part, is expanded to 6 times the original number of channels through a 1*1 convolution to obtain the fifth feature tensor of dimension 192*32*32. Subsequently, the fifth feature tensor is subjected to 3*3 depth-separable convolution processing to output the sixth feature tensor of dimension 192*32*32. Finally, a 1*1 convolution is used to reduce the number of channels of the sixth feature tensor to the number before inputting the first basic module, and finally the eighth feature tensor of dimension 32*32*32 is output.

[0020] The first subsampling layer, whose input is the eighth feature tensor of dimension 32*32*32 output by the first basic module, first processes the eighth feature tensor using a 2*2 convolutional layer with a stride of 2, and then outputs the processing result to a batch normalization layer to obtain the ninth feature tensor of dimension 64*16*16.

[0021] The second basic module takes as input the ninth feature tensor output by the first subsampling layer. First, the ninth feature vector is subjected to 1*1 convolution processing to expand the number of channels to 6 times of the input to obtain the tenth feature tensor with a dimension of 384*16*16. Then, the tenth feature tensor is subjected to 3*3 depth-separable convolution processing to output the eleventh feature tensor with a dimension of 384*16*16. Finally, a 1*1 convolution layer is used to process the feature tensor to output the twelfth feature tensor with a dimension of 64*16*16.

[0022] The second subsampling layer takes as input the twelfth feature tensor of dimension 64*16*16 output by the second basic module. It first processes the twelfth feature tensor using a 2*2 convolutional layer with a stride of 2, and then inputs the processed result into a batch normalization layer to obtain the thirteenth feature tensor of dimension 128*8*8;

[0023] The third basic module, whose input is the thirteenth feature tensor of dimension 128*8*8 output by the second subsampling layer, first performs 1*1 convolution processing on the thirteenth feature tensor to expand the number of channels to 6 times of the input to obtain a fourteenth feature tensor of dimension 768*8*8, then performs 3*3 depth-separable convolution processing on the fourteenth feature tensor to obtain a fifteenth feature tensor of dimension 768*8*8, finally, uses a 1*1 convolution layer to process the fifteenth feature tensor to output a sixteenth feature tensor of dimension 128*8*8.

[0024] The third subsampling layer, whose input is the sixteenth feature tensor with a dimension of 128*8*8 output by the third basic module, first processes the sixteenth feature tensor using a 2*2 convolutional layer with a stride of 2, and then inputs the processing result into a batch normalization layer to obtain the seventeenth feature tensor with a dimension of 256*4*4.

[0025] The fourth basic module, whose input is the seventeenth feature tensor of dimension 256*4*4 output by the third subsampling layer, first uses a 1*1 expanded convolution layer to process the seventeenth feature tensor to obtain an eighteenth feature tensor of dimension 1024*4*4, then performs a 3*3 depth-separable convolution on the eighteenth feature tensor to obtain a nineteenth feature tensor of dimension 1024*4*4, and finally, uses a 1*1 convolution to perform feature compression on the nineteenth feature tensor to output a twentieth feature tensor of dimension 256*4*4.

[0026] The last part takes the twentieth feature tensor output by the fourth basic module as input. The twentieth feature tensor is first processed using a 1*1 convolutional layer to obtain a twenty-first feature tensor with a dimension of 256*4*4. Then, the global average pooling GAP is used to convert the twenty-first feature tensor into a final feature tensor of 256*1*1. Finally, the final feature tensor is output to the fully connected classifier to obtain the final output result.

[0027] Preferably, the first detection model is obtained by distillation training through the following steps:

[0028] (2-1) 40,000 active malicious samples were obtained from VirusShare and VirusTotal websites, and 25,000 executable programs were extracted from multiple operating systems and application software as benign samples. Each benign sample and each malicious sample were converted into a binary classification image of size 128*128. All binary classification images constituted a data set, and the data set was divided into a training set and a validation set in a ratio of 8:2. The model parameters of the first detection model were initialized to obtain the initialized first detection model; specifically: the training batch size of the first detection model was set to 128, the learning rate was set to 0.001, the Adam optimization algorithm was selected, and the distillation temperature T and parameter γ were set to 1 and 1.5, respectively.

[0029] (2-2) For each sample in the training set obtained in step (2-1), input the sample into the initial part of the initialized first detection model to obtain a feature tensor corresponding to the sample with a dimension of 32*32*32;

[0030] (2-3) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-2) into the first basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 32*32*32;

[0031] (2-4) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-3) into the first subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64*16*16;

[0032] (2-5) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-4) into the second basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64*16*16;

[0033] (2-6) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-5) into the second subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128*8*8;

[0034] (2-7) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-6) into the third basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128*8*8;

[0035] (2-8) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-7) is input into the third subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256*4*4;

[0036] (2-9) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-8) into the fourth basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256*4*4;

[0037] (2-10) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the last part of the first detection model to obtain a final output feature tensor corresponding to the sample with a dimension of 256*1*1, and the final output feature tensor is input into the fully connected layer to obtain a predicted value corresponding to the sample;

[0038] (2-11) For each sample in the training set obtained in step (2-1), the total loss value L is calculated based on the predicted value corresponding to the sample obtained in step (2-10) SUM ; The total loss value is equal to:

[0039] L SUM =L GT +L MOFA ;

[0040] Among them L GT is the cross entropy loss between the predicted value corresponding to the sample obtained in step (2-10) and the true value of the predicted target in the sample, L MOFA represents the knowledge distillation loss;

[0041] (2-12) for each sample in the training set obtained in step (2-1), iteratively training the first detection model according to the total loss value of the sample obtained in step (2-11) and using a backpropagation method until the first detection model converges, thereby obtaining a preliminarily trained first detection model;

[0042] (2-13) Using the verification set obtained in step (2-1), the first detection model preliminarily trained in step (2-12) is verified to obtain the final trained first detection model.

[0043] Preferably, the knowledge distillation loss L MOFA It is obtained by the following steps:

[0044] (2-11-1) A corresponding projector is set for each stage layer of the first detection model to process features of different dimensions output by different stage layers.

[0045] (2-11-2) For each sample in the training set obtained in step (2-1), the sample is input into the trained second detection model to obtain the final output feature teacher_logits with a dimension of 512*1*1, and the predicted probability distribution P of the second detection model for the sample is calculated using the final output feature teacher_logits. t :

[0046] P t =softmax(teacher_logits / T)

[0047] (2-11-3) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-3) is input into the projector of the first stage layer to obtain the projector output feature tensor stage1_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the first stage layer is calculated through the projector output feature tensor stage1_logits OFA1 ;

[0048] (2-11-4) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-5) is input into the projector of the second stage layer to obtain a projector output feature tensor stage2_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the second stage layer is calculated based on the projector output feature tensor stage2_logits and the same method as step (2-11-3). OFA2 .

[0049] (2-11-5) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-7) is input into the projector of the third stage layer to obtain a projector output feature tensor stage3_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the third stage layer is calculated based on the projector output feature tensor stage3_logits and the same method as step (2-11-2). OFA3 .

[0050] (2-11-6) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the projector of the fourth stage layer to obtain a projector output feature tensor stage4_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the fourth stage layer is calculated based on the projector output feature tensor stage4_logits and the same method as step (2-11-3). OFA4 .

[0051] (2-11-7) For each sample in the training set obtained in step (2-1), the final output feature tensor of the dimension 256*1*1 corresponding to the sample obtained in step (2-10) is calculated, and the knowledge distillation loss L of the final output of the first detection model is calculated using the same method as step (2-11-3) OFAL

[0052] (2-11-8) The knowledge distillation loss of each stage layer obtained from step (2-11-3) to step (2-11-6) and the knowledge distillation loss of the final output of the model obtained from step (2-11-7) are weighted summed to obtain the final knowledge distillation loss L MOFA :

[0053] L MOFA =0.8*(L OFA1 +L OFA2 +L OFA3 )+0.9*L OFA4 +L OFAL ;.

[0054] Preferably, step (2-11-1) is specifically:

[0055] The first two stages of the first detection model use linear projectors. The structure of each linear projector is as follows: first, a 1*1 convolution is performed to align the number of output feature channels of the first or second stage layer of the first detection model with the output of the second detection model. Then, an adaptive average pooling layer is connected to pool the input feature map of any size into a 1*1 size. Then, a flattening layer is connected to flatten the multi-dimensional data into one dimension. Finally, a fully connected layer is connected to map the flattened features to the category output according to the feature dimension and the number of target categories.

[0056] The third stage layer uses a hybrid projector, which adds a layer of 3*3 convolution to the input head based on the above linear projector. Through 3*3 convolution and subsequent 1*1 convolution, the output feature dimension of the third stage layer of the first detection model is gradually increased to match the output feature dimension of the second detection model.

[0057] The fourth-stage layer uses a hybrid projector, which is based on the hybrid projector of the third-stage layer. A multi-head attention layer is inserted after the 1*1 two-dimensional convolution to capture the dependencies between features and strengthen the model's learning of complex relationships between features.

[0058] Steps (2-11-3) are specifically as follows:

[0059] Specifically, this step first obtains the predicted probability distribution P of the sample by the projector of the first stage layer through the projector output feature tensor stage1_logits s1 :

[0060] P s1 =softmax(stage1_logits / T)

[0061] Then, by predicting the probability distribution P s1 And the predicted probability distribution P obtained in step (2-11-2) t Calculate the knowledge distillation loss L of the first stage layer OFA1 :

[0062]

[0063] in, and Represent the predicted probability distribution P s1 and P t The sample is predicted to be its true category The probability of represents the set of all possible predicted categories, Indicates traversal of pairs except the real category All possible categories c except The cumulative sum is then calculated and the average value is taken.

[0064] Preferably, the second detection model is obtained based on the Convnextv2 model, specifically by adding a multi-head attention layer after the 7*7 convolution module therein, and adding a channel compression module after the multi-head attention layer;

[0065] The second detection model in step (3) is trained by the following steps:

[0066] A. Obtain 40,000 malicious samples from VirusShare and VirusTotal, and 25,000 executable programs from various operating systems and application software as benign samples;

[0067] B. Convert each benign sample and malicious sample obtained in step A into a binary classification image of size 128*128. All binary classification images constitute a dataset, which is divided into a training set and a validation set in a ratio of 8:2.

[0068] C. Initialize the model parameters of the second detection model to obtain an initialized second detection model. Specifically, set the training batch size of the first detection model to 128, the learning rate to 0.001, and select the Adam optimization algorithm.

[0069] D. For each sample in the training set obtained in step B (whose dimension is 3*128*128), input the sample into the initialized second detection model to output a final feature tensor corresponding to the sample with a dimension of 512*1*1, and input the final feature tensor into the fully connected layer to obtain the predicted value corresponding to the sample;

[0070] E. For each sample in the training set obtained in step B, calculate the cross entropy loss based on the predicted value corresponding to the sample obtained in step D and the true value of the sample.

[0071] F. For each sample in the training set obtained in step B, use the cross entropy loss obtained in step E and the back propagation method to iteratively train the second detection model until the second detection model converges, thereby obtaining a preliminarily trained second detection model;

[0072] G. Use the validation set obtained in step B to validate the model preliminarily trained in step F to obtain the final trained second detection model.

[0073] Preferably, in step (3), the process of updating the pre-trained second detection model to obtain an updated second detection model includes the following steps:

[0074] (3-1) Obtain all malicious code samples in the cloud malicious code data center after the data is updated in step (2).

[0075] (3-2) Analyze the updated malicious code sample obtained in step (3-1) to determine whether its attack type and attack method, as well as the malicious code family to which it belongs, are similar to the corresponding original sample in the cloud data center. If so, proceed to step (3-3); otherwise, proceed to step (3-4);

[0076] (3-3) Training the second detection model using a fine-tuning-based incremental learning method to obtain an updated second detection model; step (3-3) includes the following sub-steps:

[0077] (3-3-1) The updated malicious code samples obtained in step (3-1) are combined with the corresponding samples similar to them in the cloud malicious code data center in a ratio of 3:7 as a data set.

[0078] The dataset is divided into training set and validation set in a ratio of 8:2;

[0079] (3-3-2) Convert each sample in the data set obtained in step (3-3-1) into an RGB image. All RGB images constitute an RGB data set.

[0080] (3-3-3) Initializing the current second detection model to obtain an initialized second detection model;

[0081] (3-3-4) For each sample of the training set obtained in step (3-3-1), the sample is input into the second detection model initialized in step (3-3-3) to obtain a feature tensor with a dimension of 512*1*1, and the feature tensor is input into the fully connected layer to obtain the predicted value corresponding to the sample, and the cross entropy loss is calculated by the predicted value and the true value of the sample, and the second detection model is iteratively trained using the cross entropy loss until the second detection model converges, thereby obtaining a preliminarily trained second detection model.

[0082] (3-3-5) Use the verification set obtained in step (3-3-1) to verify the second detection model preliminarily trained in step (3-3-4) to obtain an updated second detection model.

[0083] (3-4) Training the second detection model using a regularization-based incremental learning method to obtain an updated second detection model; step (3-4) includes the following sub-steps:

[0084] (3-4-1) Using the current second detection model as the starting point for model pre-training, multiple samples randomly selected from the cloud-based malicious code data center are combined with the malicious code samples obtained in step (3-1) in a ratio of 7:3 as the sample set for incremental learning training (the purpose of which is to ensure the diversity and comprehensiveness of the training data set), and the sample set is divided into a training set and a validation set in a ratio of 8:2;

[0085] (3-4-2) Convert each sample in the sample set obtained in step (3-4-1) into an RGB image, and all RGB images constitute an RGB data set;

[0086] (3-4-3) Initialize the current second detection model to obtain the initialized second detection model, and calculate the Fisher information matrix of the initialized second detection model (used to measure the importance of each parameter in the model for data prediction).

[0087] (3-4-4) For each sample in the training set obtained in step (3-4-1), the sample is input into the second detection model initialized in step (3-4-3) to obtain the predicted value corresponding to the sample, and the cross entropy loss L is obtained based on the predicted value and the true value of the sample. GT , and calculate the total loss L based on the cross entropy loss SUM :

[0088] L SUM =L GT +∑ i F i (θ i -θ old,i ) 2 ,

[0089] Among them F i is the i-th diagonal element in the Fisher information matrix obtained in step (3-4-3), θ i is the value of the i-th parameter in the second detection model after initialization in step (3-4-3), θ old,i is the value of the i-th parameter in the second detection model before initialization in step (3-4-3), and i∈[1, the total number of all parameters in the second model];

[0090] (3-4-5) Use the total loss L obtained in step (3-4-4) SUM , and iteratively train the second detection model using the back propagation method until the second detection model converges, thereby obtaining a preliminarily trained second detection model.

[0091] (3-4-6) Use the verification set obtained in step (3-4-1) to verify the second detection model preliminarily trained in step (3-4-5) to obtain an updated second detection model.

[0092] Preferably, the process of updating the first detection model using the second detection model in step (4) includes the following sub-steps:

[0093] (4-1) Determine whether the second detection model obtained in step (3) is trained using the incremental learning method based on fine-tuning. If so, proceed to step (4-2); otherwise, proceed to step (4-3).

[0094] (4-2) Perform distillation and fine-tuning-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceed to step (4-4); step (4-2) includes the following sub-steps:

[0095] (4-2-1) Obtain the data set obtained in step (3-3-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2;

[0096] (4-2-2) Initializing the current first detection model to obtain an initialized first detection model;

[0097] Specifically, the initialization operation does not involve resetting the model weights. It only sets the learning rate to 0.0001, selects the Adam optimization algorithm, and freezes the weights of the first three stages of the first detection model before performing incremental training updates.

[0098] (4-2-3) For each sample in the training set obtained in step (4-2-1), the sample is input into the first detection model initialized in step (4-2-2) to obtain the predicted value corresponding to the sample, and the total loss L is calculated based on the predicted value. SUM1 :

[0099] L SUM1 =L GT1 +L MOFA1

[0100] Among them L GT1 Represents the cross entropy loss between the predicted value corresponding to the sample obtained in step (4-2-3) and the true value of the predicted target in the sample, L MOFA1 represents the knowledge distillation loss, and has:

[0101] L MOFA1 =0.9*L OFA4 +L OFAL

[0102] (4-2-4) Use the total loss L obtained in step (4-2-3) SUM1 , and iteratively train the first detection model using the back propagation method until the first detection model converges, and obtain the trained first detection model after verification by the verification set, and then proceed to step (4-4);

[0103] (4-3) Performing distillation and regularization-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceeding to step (4-4); step (4-3) includes the following sub-steps:

[0104] (4-3-1) Obtain the sample set obtained in step (3-4-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2.

[0105] (4-3-2) Initializing the current first detection model to obtain an initialized first detection model;

[0106] (4-3-3) For each sample in the training set obtained in step (4-3-1), the sample is input into the first detection model initialized in step (4-3-2) to extract features and calculate the total loss:

[0107] L SUM2 =L GT2 +L MOFA2 +∑ i F i (θ i -θ old,i ) 2 ,

[0108] Knowledge distillation loss L MOFA2 The value of L in the initial round of training is: MOFA2 =0.9*L OFA4 +L OFAL ,

[0109] During training, if the validation set accuracy and loss function do not decrease for 5 consecutive training rounds, the knowledge distillation loss is:

[0110] L MOFA2 =0.8L OFA3 +0.9*L OFA4 +L OFAL

[0111] (4-3-4) Use the total loss L obtained in step (4-3-3) SUM2, and use the back propagation method to iteratively train the first detection model until the first detection model converges to obtain a preliminarily trained first detection model, and use the verification set to verify the preliminarily trained first detection model to obtain a trained first detection model.

[0112] (4-4) The trained first detection model is encapsulated and serialized, and the processed first detection model is put into use at the edge.

[0113] According to another aspect of the present invention, a system for updating an IoT malicious code detection model is provided. The system is applied in an environment including a first detection model, a cloud-based malicious code data center, and a second detection model. The system includes:

[0114] The first module is set at the edge end and is used to read multiple source files from the local computer and batch convert all the source files into multiple RGB images. All RGB images constitute an image set.

[0115] The second module is set at the edge and is used to input each RGB image in the image set obtained by the first module into a pre-trained first detection model to obtain a detection result of the source file corresponding to the RGB image, and determine whether the source file contains malicious code based on the detection result. If it does contain malicious code, the source file is uploaded to the malicious code data center in the cloud for data update and transferred to the third module. Otherwise, the process ends;

[0116] The third module is set in the cloud and is used to obtain the malicious code samples updated by the cloud malicious code data center after the data of the second module is updated, and use the malicious code samples to update the pre-trained second detection model, thereby obtaining an updated second detection model;

[0117] a fourth module, which is provided in the cloud and is used to update the first detection model at the edge end using the second detection model updated by the third module to obtain an updated first detection model;

[0118] The fifth module is set in the cloud and is used to determine whether a termination instruction is received from the customer. If so, the process ends, otherwise it returns to the first module.

[0119] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0120] 1. Due to the adoption of step (2), the present invention deploys the detection model on the edge IoT device and quickly responds to malicious code in the IoT device, thereby solving the delay and real-time problems caused by the existing cloud computing method of uploading the file to be detected to the cloud and waiting for feedback.

[0121] 2. The present invention adopts steps (2-1) to (2-13), designs a lightweight detection model, and uses knowledge distillation technology to transfer the detection capabilities of the cloud-based high-performance model to the lightweight detection model, thereby improving the detection capabilities of the lightweight detection model. This solves the technical problem that the existing method based on lightweight detection models has relatively low detection accuracy due to the shallow network depth of the lightweight model, which makes it difficult to learn the deep-level features of malicious code.

[0122] 3. The present invention adopts steps (3) and (4), which first uses incremental learning technology to update the high-performance model on the cloud, and then simultaneously uses knowledge distillation technology and incremental learning technology to update the lightweight detection model. This solves the technical problem that the existing lightweight detection model-based method has poor adaptability to dynamic environments due to the difficulty of updating the detection model on edge IoT devices when the IoT environment changes rapidly, devices are constantly updated, and malicious codes are constantly evolving. BRIEF DESCRIPTION OF THE DRAWINGS

[0123] Figure 1 Schematic diagram of the application environment of the method for updating the Internet of Things malicious code detection model of the present invention;

[0124] Figure 2 is a schematic diagram of step (1) in the method of the present invention;

[0125] Figure 3 It is a structural diagram of the first detection model of the present invention;

[0126] Figure 4 is a schematic diagram of the training process of the first detection model of the present invention;

[0127] Figure 5 is a schematic structural diagram of a second detection model of the present invention;

[0128] Figure 6 It is a specific flow chart of step (3) in the present invention. DETAILED DESCRIPTION

[0129] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0130] The basic idea of ​​this invention is to make full use of the advantages of cloud computing and edge computing, and achieve efficient and real-time malicious code detection and protection through collaborative work between the cloud and the edge. Figure 1 As shown, the present invention is applied in an environment including an edge-end student detection model (hereinafter referred to as the first detection model), a cloud-based malicious code data center, and a cloud-based teacher detection model (hereinafter referred to as the second detection model). The first detection model provides real-time malicious code detection services for the edge and uploads detected malicious code to the cloud-based data center to update malicious code samples and models. The cloud-based malicious code data center is responsible for storing and analyzing malicious code samples and providing training samples to the second detection model. The second detection model trains and updates the model based on the samples provided by the data center, and obtains the updated first detection model through distillation and incremental learning, which is then downloaded to the edge.

[0131] The present invention provides a method for updating an Internet of Things malicious code detection model, comprising the following steps:

[0132] (1) The edge reads multiple source files from the local machine and converts all the source files into multiple RGB images in batches. All the RGB images constitute an image set.

[0133] like Figure 2 As shown, this step is specifically as follows: first, obtain each source file and read each byte data in the source file; then, map each three bytes of data to an RGB pixel point, where the first decimal data represents the red (R) component, the second decimal data represents the green (G) component, and the third decimal data represents the blue (B) component; then, fill the obtained multiple RGB pixels into the image to obtain an intermediate image, where the height and width of the intermediate image are determined according to the size of the source file (for example, if the binary data stream size of a file sample is 500KB, 500,000 decimal data can be obtained after conversion, then it can be generated RGB pixels, the width of the generated image can be determined as Where N = 500,000, then all RGB pixels are padded row-first, and the remaining part is padded with 0 pixels to obtain an intermediate image of size 409*409); finally, the size of the obtained intermediate image is scaled to 128*128 using bilinear interpolation to obtain the final RGB image;

[0134] (2) The edge terminal inputs each RGB image in the image set obtained in step (1) into a pre-trained first detection model to obtain a detection result of the source file corresponding to the RGB image, and determines whether the source file contains malicious code based on the detection result. If malicious code is contained, the source file is uploaded to the malicious code data center in the cloud for data update and the process proceeds to step (3). Otherwise, the process ends.

[0135] (3) obtaining, in the cloud, the malicious code samples updated by the cloud malicious code data center after the data is updated in step (2), and using the malicious code samples to update the pre-trained second detection model, thereby obtaining an updated second detection model;

[0136] (4) The cloud updates the first detection model at the edge using the second detection model updated in step (3) to obtain an updated first detection model;

[0137] (5) The cloud determines whether it has received a termination instruction from the client. If so, the process ends; otherwise, it returns to step (1);

[0138] Specifically, if Figure 3 As shown, the first detection model of the present invention is a lightweight graph convolution model, which specifically includes a starting part Stem, four basic modules, three subsampling layers and a final part. The specific structure of the model is:

[0139] The starting part Stem, whose input is an image of dimension 3*128*128, first uses a 3*3 conventional convolution with a stride of 2 to process the input image to obtain a first feature tensor of dimension 16*64*64 (the purpose is to retain spatial information while reducing the amount of computation); then, a 3*3 depth-wise separable convolution (DW-Conv) is used to process the feature tensor to obtain a second feature tensor of dimension 32*64*64 (the purpose is to capture the low-level patterns and textures of the image); subsequently, a 1*1 point-by-point convolution operation is performed on the obtained second feature tensor to obtain a third feature tensor (the purpose is to perform feature integration and increase the nonlinear capability of the network by adding a nonlinear activation function (ReLU)); finally, the third feature tensor is further sampled using a 3*3 depth-wise separable convolution with a stride of 2 to obtain a fourth feature tensor of dimension 32*32*32. In addition, the starting part Stem also adds a shortcut connection between the first layer 3*3 regular convolution output and the 1*1 point-by-point convolution output (the purpose is to help alleviate the gradient disappearance problem).

[0140] The first basic module, whose input is the fourth feature tensor of dimension 32*32*32 output by the starting part, is expanded to 6 times the original number of channels through a 1*1 convolution to obtain a fifth feature tensor of dimension 192*32*32. Subsequently, the fifth feature tensor is subjected to 3*3 depth-separable convolution processing to output a sixth feature tensor of dimension 192*32*32. Finally, a 1*1 convolution is used to reduce the number of channels of the sixth feature tensor to the number of channels of the fourth feature tensor input by the first basic module, and finally the eighth feature tensor of dimension 32*32*32 is output.

[0141] The first subsampling layer, whose input is the eighth feature tensor of dimension 32*32*32 output by the first basic module, first processes the eighth feature tensor using a 2*2 convolutional layer with a stride of 2, and then outputs the processing result to a batch normalization layer to obtain the ninth feature tensor of dimension 64*16*16 (the purpose is to stabilize the training and accelerate the convergence speed of the network through the normalization layer output).

[0142] The second basic module takes as input the ninth feature tensor output by the first subsampling layer. First, the ninth feature vector is subjected to 1*1 convolution processing to expand the number of channels to 6 times of the input to obtain the tenth feature tensor with a dimension of 384*16*16. Then, the tenth feature tensor is subjected to 3*3 depth-separable convolution processing to output the eleventh feature tensor with a dimension of 384*16*16. Finally, a 1*1 convolution layer is used to process the feature tensor to output the twelfth feature tensor with a dimension of 64*16*16.

[0143] The second subsampling layer takes as input the twelfth feature tensor of dimension 64*16*16 output by the second basic module. It first processes the twelfth feature tensor using a 2*2 convolutional layer with a stride of 2, and then inputs the processed result into a batch normalization layer to obtain the thirteenth feature tensor of dimension 128*8*8;

[0144] The third basic module, whose input is the thirteenth feature tensor of dimension 128*8*8 output by the second subsampling layer, first performs 1*1 convolution processing on the thirteenth feature tensor to expand the number of channels to 6 times of the input to obtain a fourteenth feature tensor of dimension 768*8*8, then performs 3*3 depth-separable convolution processing on the fourteenth feature tensor to obtain a fifteenth feature tensor of dimension 768*8*8, finally, uses a 1*1 convolution layer to process the fifteenth feature tensor to output a sixteenth feature tensor of dimension 128*8*8.

[0145] The third subsampling layer, whose input is the sixteenth feature tensor with a dimension of 128*8*8 output by the third basic module, first processes the sixteenth feature tensor using a 2*2 convolutional layer with a stride of 2, and then inputs the processing result into a batch normalization layer to obtain the seventeenth feature tensor with a dimension of 256*4*4.

[0146] The fourth basic module, whose input is the seventeenth feature tensor of dimension 256*4*4 output by the third subsampling layer, first uses a 1*1 expanded convolution layer to process the seventeenth feature tensor to obtain an eighteenth feature tensor of dimension 1024*4*4, then performs a 3*3 depth-separable convolution on the eighteenth feature tensor to obtain a nineteenth feature tensor of dimension 1024*4*4, and finally, uses a 1*1 convolution to perform feature compression on the nineteenth feature tensor to output a twentieth feature tensor of dimension 256*4*4.

[0147] Through the fourth basic module structure, it is possible to combine the local perception ability of the convolutional network and the global information processing ability of the Transformer at a lower computational cost to enhance network performance.

[0148] The last part (Head) takes as input the twentieth feature tensor output by the fourth basic module. First, the twentieth feature tensor is processed using a 1*1 convolutional layer to obtain a twenty-first feature tensor with a dimension of 256*4*4. Then, the twenty-first feature tensor is converted into a final feature tensor of 256*1*1 using global average pooling (GAP). Finally, the final feature tensor is output to the fully connected classifier to obtain the final output result.

[0149] The advantage of this structural design is that it can extract effective global information from the convolutional network to provide support for the final classification task.

[0150] The first detection model described in this paper incorporates efficient network modules from a deep learning architecture. To better learn from the second detection model, a similar backbone network structure is designed, with appropriate simplifications and adjustments. This design aims to enable the first detection model to effectively utilize limited edge resources while maintaining sufficient performance to perform complex malicious code detection tasks.

[0151] like Figure 4 As shown, the first detection model in the present invention is obtained by distillation training through the following steps:

[0152] (2-1) 40,000 active malicious samples are obtained from VirusShare and VirusTotal websites, and 25,000 executable programs (with a size greater than 1KB and less than 2MB) are extracted from multiple operating systems and application software as benign samples. Each benign sample and each malicious sample is converted into a binary classification image of size 128*128. All binary classification images constitute a data set, and the data set is divided into a training set and a validation set in a ratio of 8:2. The model parameters of the first detection model are initialized to obtain the initialized first detection model;

[0153] Specifically, the model parameters of the first detection model are initialized in this step as follows: the training batch size of the first detection model is set to 128, the learning rate is set to 0.001, the Adam optimization algorithm is selected, the distillation temperature T and the parameter γ are set to 1 and 1.5 respectively, and after each round of training, the accuracy, recall rate, and F1 score are output through the validation set as evaluation indicators.

[0154] (2-2) For each sample in the training set obtained in step (2-1), the sample is input into the initial part Stem of the initialized first detection model to obtain a feature tensor corresponding to the sample with a dimension of 32*32*32;

[0155] (2-3) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-2) into the first basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 32*32*32;

[0156] (2-4) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-3) into the first subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64*16*16;

[0157] (2-5) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-4) into the second basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64*16*16;

[0158] (2-6) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-5) into the second subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128*8*8;

[0159] (2-7) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-6) into the third basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128*8*8;

[0160] (2-8) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-7) is input into the third subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256*4*4;

[0161] (2-9) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-8) into the fourth basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256*4*4;

[0162] (2-10) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the last part of the first detection model to obtain a final output feature tensor corresponding to the sample with a dimension of 256*1*1, and the final output feature tensor is input into the fully connected layer to obtain a predicted value corresponding to the sample;

[0163] (2-11) For each sample in the training set obtained in step (2-1), the total loss value L is calculated based on the predicted value corresponding to the sample obtained in step (2-10) SUM ;

[0164] Specifically, the total loss value is equal to:

[0165] L SUM =L GT +L MOFA ;

[0166] Among them L GT is the cross entropy loss between the predicted value corresponding to the sample obtained in step (2-10) and the true value of the predicted target in the sample, L MOFA represents the knowledge distillation loss;

[0167] like Figure 4 As shown, the knowledge distillation loss L MOFA It is obtained by the following steps:

[0168] (2-11-1) A corresponding projector is set for each stage layer of the first detection model to process features of different dimensions output by different stage layers.

[0169] Among them, the first two stages of the first detection model are processed only by the linear projector to learn the linear feature relationship of the shallow layer of the second detection model. The structure of the linear projector is as follows: first, a 1*1 convolution is performed to align the number of output feature channels of the first or second stage layer of the first detection model with the output of the second detection model. Then, an adaptive average pooling layer is connected to pool the input feature map of any size into a 1*1 size, reduce the feature dimension and average the global information. Then, a flattening layer is connected to flatten the multi-dimensional data into one dimension, and finally a fully connected layer is connected. According to the feature dimension and the number of target categories, the flattened features are mapped to the category output; the third-stage layer uses a hybrid projector to add a layer of 3*3 convolution on the input head based on the above-mentioned linear projector, and gradually increases the output feature dimension of the third-stage layer of the first detection model to adapt to the output feature dimension of the second detection model through 3*3 convolution and subsequent 1*1 convolution; the hybrid projector of the fourth-stage layer is based on the hybrid projector of the third-stage layer, and inserts a multi-head attention layer after the 1*1 two-dimensional convolution to capture the dependency between features and enhance the model's learning of the complex relationship between features.

[0170] (2-11-2) For each sample in the training set obtained in step (2-1), the sample is input into the trained second detection model to obtain the final output feature teacher_logits with a dimension of 512*1*1, and the predicted probability distribution P of the second detection model for the sample is calculated using the final output feature teacher_logits. t :

[0171] P t =softmax(teacher_logits / T)

[0172] (2-11-3) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-3) is input into the projector of the first stage layer to obtain the projector output feature tensor stage1_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the first stage layer is calculated through the projector output feature tensor stage1_logits OFA1 ;

[0173] Specifically, this step first obtains the predicted probability distribution P of the sample by the projector of the first stage layer through the projector output feature tensor stage1_logits s1 :

[0174] P s1 =softmax(stage1_logits / T)

[0175] Then, by predicting the probability distribution P s1 And the predicted probability distribution P obtained in step (2-11-2) t Calculate the knowledge distillation loss L of the first stage layer OFA1 :

[0176]

[0177] in, and Represent the predicted probability distribution P s1 and P t The sample is predicted to be its true category The probability of represents the set of all possible predicted categories, Indicates traversal of pairs except the real category All possible categories c except The cumulative sum is then calculated and the average value is taken.

[0178] (2-11-4) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-5) is input into the projector of the second stage layer to obtain a projector output feature tensor stage2_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the second stage layer is calculated based on the projector output feature tensor stage2_logits and the same method as step (2-11-3). OFA2 .

[0179] (2-11-5) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-7) is input into the projector of the third stage layer to obtain a projector output feature tensor stage3_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the third stage layer is calculated based on the projector output feature tensor stage3_logits and the same method as step (2-11-3). OFA3 .

[0180] (2-11-6) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the projector of the fourth stage layer to obtain a projector output feature tensor stage4_logits with a dimension of 512*1*1, and the knowledge distillation loss L of the fourth stage layer is calculated based on the projector output feature tensor stage4_logits and the same method as step (2-11-3). OFA4 .

[0181] (2-11-7) For each sample in the training set obtained in step (2-1), the final output feature tensor of the dimension 256*1*1 corresponding to the sample obtained in step (2-10) is calculated, and the knowledge distillation loss L of the final output of the first detection model is calculated using the same method as step (2-11-3) OFAL

[0182] (2-11-8) The knowledge distillation loss of each stage layer obtained from step (2-11-3) to step (2-11-6) and the knowledge distillation loss of the final output of the model obtained from step (2-11-7) are weighted summed to obtain the final knowledge distillation loss L MOFA .

[0183] This step specifically uses the following formula:

[0184] L MOFA =0.8*(L OFA1 +L OFA2 +L OFA3 )+0.9*L OFA4 +L OFAL ;

[0185] The advantage of the above steps (2-11) is that by extracting feature tensors from all stage layers of the first detection model and setting the projector to match the second detection model, the knowledge of the second detection model can be learned more comprehensively, thereby improving the distillation effect.

[0186] (2-12) for each sample in the training set obtained in step (2-1), iteratively training the first detection model according to the total loss value of the sample obtained in step (2-11) and using a backpropagation method until the first detection model converges, thereby obtaining a preliminarily trained first detection model;

[0187] (2-13) using the validation set obtained in step (2-1) to validate the first detection model preliminarily trained in step (2-12) to obtain a final trained first detection model;

[0188] like Figure 5As shown, the second detection model Convnextv2-attnz of the present invention is obtained based on the Convnextv2 model, specifically by introducing a multi-head attention mechanism (Global Self-Attention) and a channel compression module (Squeeze-and-Excitation, referred to as SE) into the Convnextv2 model (specifically, a multi-head attention layer is added after the 7*7 convolution module, and a channel compression module is added after the multi-head attention layer). The specific structure of the Convnextv2 model has been described in detail in the paper "ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders" published at the 2023 Conference on Computer Vision and Pattern Recognition (CVPR) conference, and will not be repeated here.

[0189] By introducing a multi-head attention layer (Global Self-Attention) after the 7*7 large-core deep convolution on the basis of the original structure, the large-core convolution itself can expand the receptive field through its larger convolution kernel, which means that it can capture a wider range of spatial information in the input feature map. The subsequently introduced global self-attention mechanism can further process this information. By learning the dependencies between different regions, the model's global understanding of the entire input data is improved, the network's expressive power is enhanced, and the limitations of the convolution operation are compensated.

[0190] The channel compression module was then introduced to further enhance the performance of convolutional neural networks: The global self-attention mechanism considers features at all locations in the input feature map, helping the model capture long-range dependencies, understand and utilize the correlations between different regions in the image, and thus improving the model's understanding of global information. However, this global processing may cause some local but important features to be relatively neglected. The introduction of the channel compression module can supplement this. By adjusting the saliency weights of the features of each channel, this module can re-emphasize those channels that are more critical to the final task, thereby enabling the network to perform better in expressing details and local features.

[0191] The second detection model in step (3) of the present invention is obtained by training through the following steps:

[0192] A. Obtain 40,000 malicious samples from VirusShare and VirusTotal, and 25,000 executable programs (larger than 1KB and smaller than 2MB) from various operating systems and application software as benign samples;

[0193] B. Convert each benign sample and malicious sample obtained in step A into a binary classification image of size 128*128. All binary classification images constitute a dataset, which is divided into a training set and a validation set in a ratio of 8:2.

[0194] C. Initializing model parameters of the second detection model to obtain an initialized second detection model;

[0195] Specifically, the model parameters of the second detection model are initialized in this step as follows: the training batch size of the first detection model is set to 128, the learning rate is set to 0.001, and the Adam optimization algorithm is selected.

[0196] D. For each sample in the training set obtained in step B (whose dimension is 3*128*128), input the sample into the initialized second detection model to output a final feature tensor corresponding to the sample with a dimension of 512*1*1, and input the final feature tensor into the fully connected layer to obtain the predicted value corresponding to the sample;

[0197] E. For each sample in the training set obtained in step B, calculate the cross entropy loss based on the predicted value corresponding to the sample obtained in step D and the true value of the sample.

[0198] F. For each sample in the training set obtained in step B, use the cross entropy loss obtained in step E and the back propagation method to iteratively train the second detection model until the second detection model converges, thereby obtaining a preliminarily trained second detection model;

[0199] G. Use the validation set obtained in step B to validate the model preliminarily trained in step F to obtain the final trained second detection model;

[0200] like Figure 6 As shown, in step (3) of the present invention, the pre-trained second detection model is updated to obtain an updated second detection model. This process includes the following steps:

[0201] (3-1) Obtain all malicious code samples in the cloud malicious code data center after the data is updated in step (2).

[0202] (3-2) Analyze the updated malicious code sample obtained in step (3-1) to determine whether its attack type and attack method, as well as the malicious code family to which it belongs, are similar to the corresponding original sample in the cloud data center. If so, proceed to step (3-3); otherwise, proceed to step (3-4);

[0203] (3-3) training the second detection model using a fine-tuning-based incremental learning method to obtain an updated second detection model;

[0204] This step includes the following sub-steps:

[0205] (3-3-1) The updated malicious code samples obtained in step (3-1) are combined with the corresponding samples similar to them in the cloud malicious code data center in a ratio of 3:7 as a data set.

[0206] The dataset is divided into training set and validation set in a ratio of 8:2;

[0207] (3-3-2) Convert each sample in the data set obtained in step (3-3-1) into an RGB image. All RGB images constitute an RGB data set.

[0208] Specifically, the process of converting the sample into an RGB image in this step is exactly the same as the corresponding process in step (1), and will not be repeated here.

[0209] (3-3-3) Initializing the current second detection model to obtain an initialized second detection model;

[0210] Specifically, the initialization operation does not involve resetting the weights of the second detection model. It only sets the learning rate to 0.0001, selects the Adam optimization algorithm, and freezes the weights of the first three stages of the second detection model before performing incremental training updates.

[0211] (3-3-4) For each sample of the training set obtained in step (3-3-1), the sample is input into the second detection model initialized in step (3-3-3) to obtain a feature tensor with a dimension of 512*1*1, and the feature tensor is input into the fully connected layer to obtain the predicted value corresponding to the sample, and the cross entropy loss is calculated by the predicted value and the true value of the sample, and the second detection model is iteratively trained using the cross entropy loss until the second detection model converges, thereby obtaining a preliminarily trained second detection model.

[0212] (3-3-5) Use the verification set obtained in step (3-3-1) to verify the second detection model preliminarily trained in step (3-3-4) to obtain an updated second detection model.

[0213] (3-4) training the second detection model using a regularization-based incremental learning method to obtain an updated second detection model;

[0214] This step is specifically as follows:

[0215] (3-4-1) Using the current second detection model as the starting point for model pre-training, multiple samples randomly selected from the cloud-based malicious code data center are combined with the malicious code samples obtained in step (3-1) in a ratio of 7:3 as the sample set for incremental learning training (the purpose of which is to ensure the diversity and comprehensiveness of the training data set), and the sample set is divided into a training set and a validation set in a ratio of 8:2;

[0216] (3-4-2) Convert each sample in the sample set obtained in step (3-4-1) into an RGB image, and all RGB images constitute an RGB data set;

[0217] (3-4-3) Initialize the current second detection model to obtain the initialized second detection model, and calculate the Fisher information matrix of the initialized second detection model (used to measure the importance of each parameter in the model for data prediction).

[0218] Specifically, the initialization operation does not involve resetting the weights of the model, but simply sets the learning rate to 0.001 and selects the SGD stochastic gradient descent optimization algorithm to adapt to the new training stage.

[0219] (3-4-4) For each sample in the training set obtained in step (3-4-1), the sample is input into the second detection model initialized in step (3-4-3) to obtain the predicted value corresponding to the sample, and the cross entropy loss L is obtained based on the predicted value and the true value of the sample. GT , and calculate the total loss L based on the cross entropy loss SUM :

[0220] L SUM =L GT +∑ i F i (θ i -θ old,i ) 2 ,

[0221] Among them F i is the i-th diagonal element in the Fisher information matrix obtained in step (3-4-3), θ i is the value of the i-th parameter in the second detection model after initialization in step (3-4-3), θ old,i is the value of the i-th parameter in the second detection model before and after the initialization of step (3-4-3), and i∈[1, the total number of all parameters in the second model];

[0222] (3-4-5) Use the total loss L obtained in step (3-4-4) SUM, and iteratively train the second detection model using the back propagation method until the second detection model converges, thereby obtaining a preliminarily trained second detection model.

[0223] (3-4-6) Use the verification set obtained in step (3-4-1) to verify the second detection model preliminarily trained in step (3-4-5) to obtain an updated second detection model.

[0224] The advantages of the above steps (3-3) and (3-4) and their sub-steps are that they adopt a flexible update strategy and use different incremental learning methods to update the second detection model for different types of malicious code samples, thereby improving the update efficiency and effect.

[0225] Specifically, the process of using the second detection model to update the first detection model in step (4) above includes the following sub-steps:

[0226] (4-1) Determine whether the second detection model obtained in step (3) is trained using the incremental learning method based on fine-tuning. If so, proceed to step (4-2); otherwise, proceed to step (4-3).

[0227] (4-2) Performing distillation and fine-tuning-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceeding to step (4-4);

[0228] This step includes the following sub-steps:

[0229] (4-2-1) Obtain the data set obtained in step (3-3-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2;

[0230] (4-2-2) Initializing the current first detection model to obtain an initialized first detection model;

[0231] Specifically, the initialization operation does not involve resetting the model weights. It only sets the learning rate to 0.0001, selects the Adam optimization algorithm, and freezes the weights of the first three stages of the first detection model before performing incremental training updates.

[0232] (4-2-3) For each sample in the training set obtained in step (4-2-1), the sample is input into the first detection model initialized in step (4-2-2) to obtain the predicted value corresponding to the sample, and the total loss L is calculated based on the predicted value. SUM1 :

[0233] L SUM1 =L GT1 +L MOFA1

[0234] Among them L GT1 Represents the cross entropy loss between the predicted value corresponding to the sample obtained in step (4-2-3) and the true value of the predicted target in the sample, L MOFA1 represents the knowledge distillation loss;

[0235] The loss function in this step is obtained using the same method as in step (2-11). However, it should be noted that since the parameter weights of the first three stages of the model are frozen before training, there is no additional value in calculating the knowledge distillation loss of these stages during the distillation process. Therefore, the loss function is used in calculating the final knowledge distillation loss L. MOFA When , we only need to calculate the knowledge distillation loss between the fourth stage layer and the final output of the model, that is:

[0236] L MOFA1 =0.9*L OFA4 +L OFAL

[0237] (4-2-4) Use the total loss L obtained in step (4-2-3) SUM1 , and iteratively train the first detection model using the back propagation method until the first detection model converges, and obtain the trained first detection model after verification by the verification set, and then proceed to step (4-4);

[0238] (4-3) performing distillation and regularization-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceeding to step (4-4);

[0239] This step includes the following sub-steps:

[0240] (4-3-1) Obtain the sample set obtained in step (3-4-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2.

[0241] (4-3-2) Initializing the current first detection model to obtain an initialized first detection model;

[0242] Specifically, the initialization operation does not involve resetting the model weights, but only sets the learning rate to 0.001, selects the Sgd optimization algorithm, and freezes the weights of the first three stages of the first detection model before performing incremental training updates;

[0243] (4-3-3) For each sample in the training set obtained in step (4-3-1), the sample is input into the first detection model initialized in step (4-3-2) to extract features and calculate the total loss:

[0244] L SUM2=L GT2 +L MOFA2 +∑ i F i (θ i -θ old,i ) 2 ,

[0245] The last term in the formula is ∑ i F i (θ i -θ old,i ) 2 The calculation method is exactly the same as in the previous article and will not be repeated here. MOFA2 The value of during the initial round of training is: L MOFA2 =0.9*L OFA4 +L OFAL ,

[0246] During the training process, if the validation set accuracy and loss function do not decrease significantly after 5 consecutive training rounds, the third stage layer of the model is unfrozen. At this time, the knowledge distillation loss is:

[0247] L MOFA2 =0.8L OFA3 +0.9*L OFA4 +L OFAL ,

[0248] The same logic is applied later, and the second-stage layers and the first-stage layers of the first detection model are gradually unfrozen. Correspondingly, during the training process after unfreezing, the total knowledge distillation loss L MOFA2 The calculation of should also add the knowledge distillation loss of the unfrozen stage layer multiplied by the corresponding weight.

[0249] (4-3-4) Use the total loss L obtained in step (4-3-3) SUM2 , and use the back propagation method to iteratively train the first detection model until the first detection model converges to obtain a preliminarily trained first detection model, and use the verification set to verify the preliminarily trained first detection model to obtain a trained first detection model.

[0250] (4-4) The trained first detection model is encapsulated and serialized, and the processed first detection model is put into use at the edge.

[0251] Through these steps, adopting different incremental + distillation collaborative update strategies for the first detection model can ensure that the new knowledge acquired by the second detection model on the cloud through incremental learning can be efficiently and effectively transferred to the first detection model.

[0252] Based on the above technical solution, it can be understood that the present application proposes a method and system for updating an IoT malicious code detection model. Through the distillation method, the key knowledge and behavior of the complex second detection model are transferred to the lighter first detection model, thereby improving the detection capability of the edge first detection model. The lightweight design of the first detection model enables a rapid response to malicious activities at the edge, reduces the need for data transmission to the cloud, and reduces dependence on cloud computing resources. In addition, by uploading and updating the cloud malicious code database through edge detection to update the model, the first detection model can be continuously updated to adapt to new or changing malicious code threats.

[0253] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for updating an IoT malicious code detection model, applied in an environment including a first detection model, a cloud-based malicious code data center, and a second detection model. The first detection model is a lightweight graph convolutional model. The second detection model is based on the Convnextv2 model and is specifically obtained by adding a multi-head attention layer after the 7×7 convolution module and a channel compression module after the multi-head attention layer. The method is characterized in that: The updating method comprises the following steps: (1) The edge reads multiple source files from the local computer and converts all source files into multiple RGB images in batches. All RGB images constitute an image set. (2) The edge inputs each RGB image in the image set obtained in step (1) into the pre-trained first detection model to obtain the detection result of the source file corresponding to the RGB image, and determines whether the source file contains malicious code based on the detection result. If it contains malicious code, the source file is uploaded to the malicious code data center in the cloud for data update and proceeds to step (3). Otherwise, the process ends; (3) Obtaining in the cloud the updated malicious code sample from the cloud malicious code data center after the data update in step (2), and using the malicious code sample to update the pre-trained second detection model, thereby obtaining an updated second detection model; step (3) includes the following steps: (3-1) Obtaining the malicious code samples updated by the cloud-based malicious code data center after the data is updated in step (2); (3-2) Analyze the updated malicious code sample obtained in step (3-1) to determine whether its attack type and attack method, as well as the malicious code family to which it belongs, are similar to the corresponding original sample in the cloud data center. If so, proceed to step (3-3); otherwise, proceed to step (3-4); (3-3) training the second detection model using a fine-tuning-based incremental learning method to obtain an updated second detection model; (3-4) training the second detection model using a regularization-based incremental learning method to obtain an updated second detection model; (4) The cloud updates the first detection model at the edge using the second detection model updated in step (3) to obtain an updated first detection model; step (4) includes the following sub-steps: (4-1) Determine whether the second detection model obtained in step (3) is trained using the incremental learning method based on fine-tuning. If so, proceed to step (4-2); otherwise, proceed to step (4-3); (4-2) Perform distillation and fine-tuning-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceed to step (4-4); (4-3) performing distillation and regularization-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceeding to step (4-4); (4-4) Encapsulating and serializing the updated and trained first detection model, and putting the processed first detection model into use at the edge; (5) The cloud determines whether it has received a termination instruction from the client. If so, the process ends; otherwise, it returns to step (1).

2. The method for updating the Internet of Things malicious code detection model according to claim 1, characterized in that: Step (1) is specifically as follows: first, obtain each source file and read each byte data in the source file; then, map every three bytes of data to an RGB pixel point, where the first byte data represents the red component, the second byte data represents the green component, and the third byte data represents the blue component; then, fill the obtained multiple RGB pixel points into the image to obtain an intermediate image, the height and width of which are determined according to the size of the source file; finally, use bilinear interpolation to scale the size of the obtained intermediate image to 128×128 to obtain the final RGB image.

3. The method for updating the Internet of Things malicious code detection model according to claim 1 or 2, characterized in that: The first detection model specifically includes the starting part, four basic modules, three subsampling layers and the final part. The specific structure of the model is: In the starting part, the input is an image of dimension 3×128×128. First, a 3×3 regular convolution with a stride of 2 is used to process the input image to obtain a first feature tensor of dimension 16×64×64; then, a 3×3 depthwise separable convolution DW-Conv is used to process the feature tensor to obtain a second feature tensor of dimension 32×64×64; then, a 1×1 pointwise convolution operation is performed on the obtained second feature tensor to obtain a third feature tensor; finally, the third feature tensor is further sampled using a 3×3 depthwise separable convolution with a stride of 2 to obtain a fourth feature tensor of dimension 32×32×32; in addition, the starting part also adds a shortcut connection between the output of the first layer 3×3 regular convolution and the output of the 1×1 pointwise convolution; The first basic module, whose input is the fourth feature tensor of dimension 32×32×32 output by the starting part, is expanded to 6 times the original number of channels through a 1×1 convolution to obtain a fifth feature tensor of dimension 192×32×32. Subsequently, the fifth feature tensor is subjected to a 3×3 depthwise separable convolution to output a sixth feature tensor of dimension 192×32×32. Finally, a 1×1 convolution is used to reduce the number of channels of the sixth feature tensor to the number of channels of the fourth feature tensor input by the first basic module, and finally the eighth feature tensor of dimension 32×32×32 is output; The first subsampling layer takes as input the eighth feature tensor of dimension 32×32×32 output by the first basic module. It first processes the eighth feature tensor using a 2×2 convolutional layer with a stride of 2, and then outputs the processed result to a batch normalization layer to obtain the ninth feature tensor of dimension 64×16×16; The second basic module takes as input the ninth feature tensor output by the first subsampling layer. It first performs a 1×1 convolution on the ninth feature tensor to expand the number of channels to 6 times the input to obtain a tenth feature tensor with a dimension of 384×16×16. Then, it performs a 3×3 depthwise separable convolution on the tenth feature tensor to output an eleventh feature tensor with a dimension of 384×16×16. Finally, it uses a 1×1 convolution layer to process the feature tensor and output a twelfth feature tensor with a dimension of 64×16×16. The second subsampling layer takes as input the 12th feature tensor of dimension 64×16×16 output by the second basic module. It first processes the 12th feature tensor using a 2×2 convolutional layer with a stride of 2. The result is then fed into a batch normalization layer to obtain the 13th feature tensor of dimension 128×8×8. The third basic module, whose input is the thirteenth feature tensor of dimension 128×8×8 output by the second subsampling layer, first performs a 1×1 convolution on the thirteenth feature tensor to expand the number of channels to 6 times the input to obtain a fourteenth feature tensor of dimension 768×8×8, then performs a 3×3 depthwise separable convolution on the fourteenth feature tensor to obtain a fifteenth feature tensor of dimension 768×8×8, and finally, uses a 1×1 convolution layer to process the fifteenth feature tensor to output a sixteenth feature tensor of dimension 128×8×8; The third subsampling layer takes as input the 16th feature tensor of dimension 128×8×8 output by the third basic module. It first processes the 16th feature tensor using a 2×2 convolutional layer with a stride of 2. The result is then fed into a batch normalization layer to obtain the 17th feature tensor of dimension 256×4×4. The fourth basic module, whose input is the 17th feature tensor with a dimension of 256×4×4 output by the third subsampling layer, first processes the 17th feature tensor using a 1×1 expanded convolution layer to obtain a 1024-dimensional feature tensor. The eighteenth feature tensor is then subjected to a 3×3 depthwise separable convolution to obtain a 1024-dimensional feature tensor. Finally, a 1×1 convolution is used to perform feature compression on the nineteenth feature tensor to output a twentieth feature tensor with a dimension of 256×4×4; The last part, whose input is the twentieth feature tensor output by the fourth basic module, is first processed using a 1×1 convolution layer to obtain a 256× Then, global average pooling GAP is used to convert the 21st feature tensor into a final feature tensor of 256×1×1. Finally, the final feature tensor is output to the fully connected classifier to obtain the final output result.

4. The method for updating the Internet of Things malicious code detection model according to claim 3, characterized in that: The first detection model is trained by distillation through the following steps: (2-1) 40,000 active malicious samples were obtained from VirusShare and VirusTotal websites, and 25,000 executable programs were extracted from multiple operating systems and application software as benign samples. Each benign sample and each malicious sample were converted into a binary classification image of size 128×128. All binary classification images constituted a data set, which was divided into a training set and a validation set in a ratio of 8:

2. The model parameters of the first detection model were initialized to obtain the initialized first detection model; specifically, the training batch size of the first detection model was set to 128, the learning rate was set to 0.001, the Adam optimization algorithm was selected, and the distillation temperature T and the parameters were set to 0. Set to 1 and 1.5 respectively; (2-2) For each sample in the training set obtained in step (2-1), input the sample into the initial part of the initialized first detection model to obtain a feature tensor of dimension 32×32×32 corresponding to the sample; (2-3) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-2) into the first basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 32×32×32; (2-4) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-3) into the first subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64×16×16; (2-5) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-4) into the second basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 64×16×16; (2-6) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-5) into the second subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128×8×8; (2-7) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-6) into the third basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 128×8×8; (2-8) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-7) into the third subsampling layer of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256×4×4; (2-9) For each sample in the training set obtained in step (2-1), input the feature tensor corresponding to the sample obtained in step (2-8) into the fourth basic module of the first detection model to obtain a feature tensor corresponding to the sample with a dimension of 256×4×4; (2-10) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the last part of the first detection model to obtain a final output feature tensor corresponding to the sample with a dimension of 256×1×1, and the final output feature tensor is input into the fully connected layer to obtain a prediction value corresponding to the sample; (2-11) For each sample in the training set obtained in step (2-1), calculate the total loss value based on the predicted value corresponding to the sample obtained in step (2-10) ; The total loss value is equal to: ; in The cross entropy loss between the predicted value corresponding to the sample obtained in step (2-10) and the true value of the predicted target in the sample, represents the knowledge distillation loss; Knowledge Distillation Loss It is obtained by the following steps: (2-11-1) Set a corresponding projector for each stage layer of the first detection model to process features of different dimensions output by different stage layers; (2-11-2) For each sample in the training set obtained in step (2-1), the sample is input into the trained second detection model to obtain the final output feature teacher_logits with a dimension of 512×1×1, and the predicted probability distribution of the second detection model for the sample is calculated using the final output feature teacher_logits. : ; (2-11-3) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-3) is input into the projector of the first stage layer to obtain the projector output feature tensor stage1_logits with a dimension of 512×1×1, and the knowledge distillation loss of the first stage layer is calculated through the projector output feature tensor stage1_logits ; (2-11-4) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-5) is input into the projector of the second stage layer to obtain the projector output feature tensor stage2_logits with a dimension of 512×1×1, and the second stage layer is calculated based on the projector output feature tensor stage2_logits and the same method as step (2-11-3). ; (2-11-5) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-7) is input into the projector of the third stage layer to obtain the projector output feature tensor stage3_logits with a dimension of 512×1×1, and the third stage layer is calculated based on the projector output feature tensor stage3_logits and the same method as step (2-11-3). ; (2-11-6) For each sample in the training set obtained in step (2-1), the feature tensor corresponding to the sample obtained in step (2-9) is input into the projector of the fourth stage layer to obtain the projector output feature tensor stage4_logits with a dimension of 512×1×1, and the fourth stage layer is calculated based on the projector output feature tensor stage4_logits and the same method as step (2-11-3). ; (2-11-7) For each sample in the training set obtained in step (2-1), the final output feature tensor of the dimension 256×1×1 corresponding to the sample obtained in step (2-10) is calculated using the same method as in step (2-11-3) and the final output of the first detection model is calculated. (2-11-8) The layers of each stage obtained from step (2-11-3) to step (2-11-6) , and the final output of the model obtained in step (2-11-7) Sum the weights to get the final : ; (2-12) For each sample in the training set obtained in step (2-1), iteratively train the first detection model using the backpropagation method according to the total loss value of the sample obtained in step (2-11) until the first detection model converges, thereby obtaining a preliminarily trained first detection model; (2-13) Using the validation set obtained in step (2-1), the first detection model preliminarily trained in step (2-12) is validated to obtain the final trained first detection model.

5. The method for updating the Internet of Things malicious code detection model according to claim 4, characterized in that: Step (2-11-1) is specifically: The first two stages of the first detection model use linear projectors. Each linear projector has the following structure: first, a 1×1 convolution to align the number of channels of the output feature maps of the first or second stage layer of the first detection model with the output of the second detection model. Then, an adaptive average pooling layer is used to pool the input feature map of any size into a 1×1 size. Then, a flattening layer is used to flatten the multidimensional data into one dimension. Finally, a fully connected layer is used to map the flattened features to the category output according to the feature dimension and the number of target categories. The third stage layer uses a hybrid projector, which adds a layer of 3×3 convolution to the input head based on the above linear projector. Through 3×3 convolution and subsequent 1×1 convolution, the output feature dimension of the third stage layer of the first detection model is gradually increased to match the output feature dimension of the second detection model. The fourth stage layer uses a hybrid projector. Based on the hybrid projector of the third stage layer, a multi-head attention layer is inserted after the 1×1 two-dimensional convolution to capture the dependencies between features and strengthen the model's learning of complex relationships between features. The specific steps (2-11-3) are: Specifically, this step first obtains the predicted probability distribution of the sample by the projector of the first stage layer through the projector output feature tensor stage1_logits : ; Then, by predicting the probability distribution And the predicted probability distribution obtained in step (2-11-2) Calculate the knowledge distillation loss of the first stage layer : ; in, and Represent the predicted probability distribution and The sample is predicted to be its true category The probability of represents the set of all possible predicted categories, Indicates traversal of pairs except the real category All possible categories c except The cumulative sum is then calculated and the average value is taken.

6. The method for updating the Internet of Things malicious code detection model according to claim 5, characterized in that: The second detection model in step (3) is trained by the following steps: A. Obtain 40,000 malicious samples from VirusShare and VirusTotal, and 25,000 executable programs from various operating systems and application software as benign samples; B. Convert each benign and malicious sample obtained in step A into a binary image of size 128×128. All binary images constitute a dataset, which is then divided into a training set and a validation set in a ratio of 8:

2. C. Initialize the model parameters of the second detection model to obtain an initialized second detection model; specifically, set the training batch size of the first detection model to 128, the learning rate to 0.001, and select the Adam optimization algorithm; D. For each sample in the training set obtained in step B, input the sample into the initialized second detection model to output a final feature tensor of dimension 512×1×1 corresponding to the sample, and input the final feature tensor into the fully connected layer to obtain the prediction value corresponding to the sample; E. For each sample in the training set obtained in step B, calculate the cross entropy loss based on the predicted value corresponding to the sample obtained in step D and the true value of the sample; F. For each sample in the training set obtained in step B, use the cross entropy loss obtained in step E and the back propagation method to iteratively train the second detection model until the second detection model converges, thereby obtaining a preliminarily trained second detection model; G. Use the validation set obtained in step B to validate the model preliminarily trained in step F to obtain the final trained second detection model.

7. The method for updating the Internet of Things malicious code detection model according to claim 6, characterized in that: Step (3-3) includes the following sub-steps: (3-3-1) The updated malicious code samples obtained in step (3-1) are combined with similar corresponding samples in the cloud malicious code data center in a ratio of 3:7 as a dataset, and the dataset is divided into a training set and a validation set in a ratio of 8:2; (3-3-2) Convert each sample in the data set obtained in step (3-3-1) into an RGB image. All RGB images constitute the RGB data set. (3-3-3) Initializing the current second detection model to obtain an initialized second detection model; (3-3-4) For each sample in the training set obtained in step (3-3-1), the sample is input into the second detection model initialized in step (3-3-3) to obtain a feature tensor with a dimension of 512×1×1, and the feature tensor is input into the fully connected layer to obtain the predicted value corresponding to the sample, and the cross entropy loss is calculated by the predicted value and the true value of the sample. The second detection model is iteratively trained using the cross entropy loss until the second detection model converges, thereby obtaining a preliminarily trained second detection model; (3-3-5) Using the validation set obtained in step (3-3-1), the second detection model preliminarily trained in step (3-3-4) is validated to obtain an updated second detection model; step (3-4) includes the following sub-steps: (3-4-1) Using the current second detection model as the starting point for model pre-training, multiple samples randomly selected from the cloud-based malicious code data center are combined with the malicious code samples obtained in step (3-1) in a ratio of 7:3 as the sample set for incremental learning training, and the sample set is divided into a training set and a validation set in a ratio of 8:2; (3-4-2) Convert each sample in the sample set obtained in step (3-4-1) into an RGB image. All RGB images constitute an RGB data set. (3-4-3) Initializing the current second detection model to obtain an initialized second detection model, and calculating the Fisher information matrix of the initialized second detection model; (3-4-4) For each sample in the training set obtained in step (3-4-1), the sample is input into the second detection model initialized in step (3-4-3) to obtain the predicted value corresponding to the sample, and the cross entropy loss is obtained based on the predicted value and the true value of the sample. , and calculate the total loss based on this cross entropy loss : , in is the i-th diagonal element in the Fisher information matrix obtained in step (3-4-3), is the value of the i-th parameter in the second detection model after initialization in step (3-4-3), is the value of the i-th parameter in the second detection model before initialization in step (3-4-3), and i∈[1, the total number of all parameters in the second model]; (3-4-5) Total loss obtained using step (3-4-4) , and iteratively train the second detection model using a back propagation method until the second detection model converges, thereby obtaining a preliminarily trained second detection model; (3-4-6) Use the validation set obtained in step (3-4-1) to validate the second detection model preliminarily trained in step (3-4-5) to obtain an updated second detection model.

8. The method for updating the Internet of Things malicious code detection model according to claim 7, characterized in that: Step (4-2) includes the following sub-steps: (4-2-1) Obtain the data set obtained in step (3-3-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2; (4-2-2) Initializing the current first detection model to obtain an initialized first detection model; (4-2-3) For each sample in the training set obtained in step (4-2-1), input the sample into the first detection model initialized in step (4-2-2) to obtain the predicted value corresponding to the sample, and calculate the total loss based on the predicted value : ; in Represents the cross entropy loss between the predicted value corresponding to the sample obtained in step (4-2-3) and the true value of the predicted target in the sample, represents the knowledge distillation loss, and has: ; (4-2-4) Total loss obtained using step (4-2-3) , and iteratively train the first detection model using the back propagation method until the first detection model converges, and obtain the trained first detection model after verification by the verification set, and then proceed to step (4-4); Step (4-3) includes the following sub-steps: (4-3-1) Obtain the sample set obtained in step (3-4-1) as the data set for training the first detection model, and divide the data set into a training set and a validation set in a ratio of 8:2; (4-3-2) Initializing the current first detection model to obtain an initialized first detection model; (4-3-3) For each sample in the training set obtained in step (4-3-1), input the sample into the first detection model initialized in step (4-3-2) to extract features and calculate the total loss: ; Knowledge Distillation Loss The value of during the initial round of training is: , During training, if the validation set accuracy and loss function do not decrease for 5 consecutive training rounds, the knowledge distillation loss is: ; (4-3-4) Total loss obtained using step (4-3-3) , and use the back propagation method to iteratively train the first detection model until the first detection model converges to obtain a preliminarily trained first detection model, and use the verification set to verify the preliminarily trained first detection model to obtain a trained first detection model.

9. A system for updating an IoT malicious code detection model, applied in an environment including a first detection model, a cloud-based malicious code data center, and a second detection model. The first detection model is a lightweight graph convolutional model, and the second detection model is based on the Convnextv2 model. Specifically, the second detection model is obtained by adding a multi-head attention layer after the 7×7 convolution module and a channel compression module after the multi-head attention layer. The system is characterized by: The updating system comprises: The first module is set at the edge end and is used to read multiple source files from the local computer and batch convert all the source files into multiple RGB images. All RGB images constitute an image set. The second module is set at the edge and is used to input each RGB image in the image set obtained by the first module into a pre-trained first detection model to obtain a detection result of the source file corresponding to the RGB image, and determine whether the source file contains malicious code based on the detection result. If it does contain malicious code, the source file is uploaded to the malicious code data center in the cloud for data update and transferred to the third module. Otherwise, the process ends; The third module is set in the cloud and is used to obtain the malicious code samples updated by the cloud malicious code data center after the data of the second module is updated, and use the malicious code samples to update the pre-trained second detection model, thereby obtaining an updated second detection model; this process includes the following steps: (3-1) Obtaining the malicious code samples updated by the cloud-based malicious code data center after the second module data is updated; (3-2) Analyze the updated malicious code sample obtained in step (3-1) to determine whether its attack type and attack method, as well as the malicious code family to which it belongs, are similar to the corresponding original sample in the cloud data center. If so, proceed to step (3-3); otherwise, proceed to step (3-4); (3-3) training the second detection model using a fine-tuning-based incremental learning method to obtain an updated second detection model; (3-4) training the second detection model using a regularization-based incremental learning method to obtain an updated second detection model; The fourth module is set in the cloud and is used to update the first detection model at the edge using the second detection model updated by the third module to obtain an updated first detection model. This process includes the following sub-steps: (4-1) Determine whether the second detection model obtained by the third module is trained using the incremental learning method based on fine-tuning. If so, proceed to step (4-2); otherwise, proceed to step (4-3); (4-2) Perform distillation and fine-tuning-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceed to step (4-4); (4-3) performing distillation and regularization-based incremental learning on the first detection model simultaneously to obtain an updated first detection model, and then proceeding to step (4-4); (4-4) Encapsulating and serializing the updated and trained first detection model, and putting the processed first detection model into use at the edge; The fifth module is set in the cloud and is used to determine whether a termination instruction is received from the customer. If so, the process ends, otherwise it returns to the first module.

Citation Information

Patent Citations

  • Internet of Things malicious software family classification method based on lightweight convolutional neural network and multi-teacher knowledge distillation

    CN116541837A

  • Automated feature extraction and artificial intelligence (AI) based detection and classification of malware

    US20200045063A1