A Malware Detection Method and System Based on an Adaptive Computing Time Strategy

By introducing DACT-Transformer, a deep learning framework with adaptive computing time strategies in malware detection, the problems of performance limitations and high computing load in the prior art are solved, and more efficient malware detection is achieved.

CN115408693BActive Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211071188.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-06-13
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

The existing malware detection methods have performance limitations and require high computing load on the device, making it difficult to promote and apply in practice.

Method used

A deep learning framework DACT-Transformer based on adaptive computing time strategy is proposed to detect Android malware. The framework reduces computational load and reduces equipment pressure by controlling the number of Transformer blocks.

Benefits of technology

With slightly reduced accuracy, DACT-Transformer significantly reduces the computational load of the model, reduces the pressure on the device, and improves the overall performance of malware detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408693B_ABST
    Figure CN115408693B_ABST
Patent Text Reader

Abstract

The present invention discloses a malicious software detection method and system based on an adaptive computing time strategy. The present invention proposes a new deep learning framework - DACT-Transformer for detecting Android malicious software, which is an adaptive computing time strategy for a Transformer-like model. DACT-Transformer adds an adaptive computing mechanism to the conventional processing of Transformer, which can control the number of Transformer blocks to be executed during inference. By doing so, the present invention reduces the computational load of the model while slightly reducing the accuracy, and reduces the pressure on the device in actual production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software testing, and particularly relates to a malicious software detection method and system based on an adaptive computing time strategy. Background Art

[0002] Malicious software refers to any software that causes damage to a computer system. They can steal protected data, delete documents, or add software without user approval. Different malicious software, such as adware, banking malware, SMS malware, etc., are constantly updated with the rapid development of encryption and obfuscation technologies, becoming one of the serious threats to the cyber space. Therefore, it is crucial for user protection to timely identify whether the downloaded software sample is legal.

[0003] The current main malicious software detection methods mainly include the malicious software detection method based on signature, the malicious software detection method based on machine learning, the malicious software detection method based on deep learning, etc.:

[0004] 1) The basic principle of the malicious software detection method based on signature is to use the unique feature information of each software for matching. That is, when the signature of a known malicious software is specified, by matching it with the signature of the target software to be detected, if the same signature is found in the existing malicious software signature database, the target software is determined to be malicious software, otherwise it is benign software. For this type of method, the existing malicious software signature database is the key. Therefore, the disadvantage of this type of method is also that it is easy for attackers to change the software syntax structure through obfuscation technology, thereby changing the signature. And although it can achieve the purpose of malicious detection through exact matching with the existing signatures, it cannot handle unknown malicious software.

[0005] 2) The proposal of the malicious software detection method based on machine learning solves the above problems. The basic principle of the malicious software detection method based on machine learning is to extract different feature descriptions of different behaviors of the sample to be analyzed through techniques such as program analysis, and then each sample is represented by a fixed-dimensional vector. Finally, with the help of existing machine learning algorithms, the samples with known labels are trained to construct a classifier, so as to be able to predict and judge unknown samples. However, the problem with the malicious software detection method based on machine learning is that most of the current work relies on a large number of samples for supervised learning to train a better model.

[0006] 3) Deep learning is a method in machine learning based on representation learning of data. Deep learning technology can integrate feature learning into the model building process, thereby reducing the incompleteness caused by artificially designed features and avoiding the defect that machine learning depends on a large number of samples. Although the existing research work related to malware detection has made good progress, there are still great limitations in the performance of existing methods. The latest technologies, while improving performance, also increase the requirements for device computing load, which makes the existing models still in the theoretical research stage and difficult to be popularized and applied in practice. Therefore, new models that can reduce the device computing load are urgently needed to be studied. Summary of the Invention

[0007] The first object of the present invention is to propose a new deep learning framework - DACT-Transformer for detecting Android malware in view of the deficiencies of the prior art. It is an adaptive computing time strategy for models similar to the Transformer. DACT-Transformer adds an adaptive computing mechanism to the conventional processing of the Transformer, which can control the number of Transformer blocks to be executed during inference. By doing so, the present invention reduces the computing load of the model while slightly reducing the accuracy, and reduces the pressure on the device in actual production.

[0008] The present invention includes the following steps:

[0009] S1. Collect the APK data of current Android software, with labels of malware and benign software; clean the APK data and divide the data to obtain a training set, a validation set, and a test set; the APK data refers to decompressing the APK installation package and parsing the decompression result to extract multiple features of the program, and the multiple features include system calls, binders, and the call frequencies of composite behaviors.

[0010] S2. Use the convolutional neural network Conv1d to expand the two-dimensional data into a multi-dimensional dense vector for the data processed in step S1, embed the weak correlation information between the original features into the continuous vector space, and generate a sequence cluster matrix.

[0011] S3. Build a malware detection model - DACT-Transformer based on the adaptive computing time strategy, and use the training set data to train the model;

[0012] S4. Use the trained DACT-Transformer model to detect the test set data.

[0013] S5. Validate the validation set data using the tested DACT-Transformer model.

[0014] S6. Use the validated DACT-Transformer model to detect Android software.

[0015] Furthermore, the method of data cleaning in step S1 is as follows:

[0016] Clean missing values and outliers, quantitatively encode the feature attributes of the data, and normalize the data.

[0017] Even further, the method of data cleaning in step S1 is as follows:

[0018] 1) Check if there are missing values. If there are, continue to determine if the number of missing values is greater than the preset threshold A. If so, it is considered that there are too many missing values, and the average or mode of the sample set attributes with the same sample label is selected to fill the missing values. If not, it is considered that there are few missing values, and the data row where the missing value is located can be deleted;

[0019] 2) Check if there are outliers. If there are, continue to determine if the number of outliers is greater than the preset threshold B. If so, it is considered that there are too many outliers, and the cause of the outliers needs to be analyzed and the outliers modified. If not, it is considered that there are few outliers, and the data row including the outliers is deleted;

[0020] 3) Feature attribute quantization encoding: Encode the discrete feature attributes within the range 0 - m, where m + 1 represents the total number of software types.

[0021] 4) The normalization method adopted is to impose the following constraints on each feature. The formula is as follows:

[0022]

[0023] where, X i,norm represents the normalization result of the i-th feature X i , and X i,min , X i,max represent the minimum and maximum values of the column where the i-th feature is located (i.e., all features of the same type as X i ) respectively.

[0024] Furthermore, in step S2, the model is trained using the training data to obtain the final detection model.

[0025] Furthermore, the steps of using the detection model for testing in step S3 are as follows: After cleaning the test data samples, input them into the DACT-Transformer model for detection and output the detection results.

[0026] The DACT-Transformer model includes an Embedding layer, N Transformer encoding layers, and N adaptive computing layers DACT; N = 12;

[0027] The Embedding layer is used to expand the input sequence into a feature vector of 512 dimensions through Embedding;

[0028] The first Transformer encoding layer receives the output of the Embedding layer. After being encoded by the Transformer, the output is copied into two parts, which are respectively input into the second Transformer encoding layer and the first DACT layer.

[0029] The nth Transformer encoding layer receives the output of the previous Transformer encoding layer. After being encoded by the Transformer, the output is copied into two parts, which are respectively input into the (n + 1)th Transformer encoding layer and the nth DACT layer; where 2 ≤ n ≤ N - 1;

[0030] The Nth Transformer encoding layer receives the output of the previous Transformer encoding layer. After being encoded by the Transformer, the output is input into the Nth DACT layer;

[0031] The first DACT layer receives the output of the first Transformer encoding layer. After processing the input matrix through Softmax and Sigmoid, it obtains an intermediate prediction result y n and its confidence, and then outputs them to the second DACT layer;

[0032] The nth DACT layer receives the outputs of the previous Transformer encoding layer and DACT layer. After processing the input matrix through Softmax and Sigmoid, it obtains an intermediate prediction result y n and its confidence, and then outputs them to the next DACT layer; In each DACT layer, when the confidence reaches a certain range interval, the training is terminated, and the current prediction result is used as the intermediate prediction result for output; 2 ≤ n ≤ N - 1;

[0033] The Nth DACT layer receives the outputs of the Nth Transformer encoding layer and the (N - 1)th DACT layer. After processing the input matrix through Softmax and Sigmoid, it obtains the final prediction result y N 。

[0034] The Transformer encoding layer includes a MultiHead Attention layer and a fully-connected layer. The MultiHead Attention layer is used to calculate the attention distribution of the input feature vectors through Self-Attention to obtain a matrix. The fully-connected layer is used to perform a non-linear transformation on the features after the output of the MultiHead Attention layer is normalized.

[0035] The second object of the present invention is to provide a malware detection system for implementing the method, which is characterized by including a malware detection model DACT-Transformer based on an adaptive computing time strategy that has been trained, verified, and tested.

[0036] The third object of the present invention is to provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method.

[0037] The fourth object of the present invention is to provide a computing device including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method is implemented. Compared with the prior art, the beneficial effects of the present invention are:

[0038] The malware detection method based on Transformer of the present invention uses a method that combines the large-scale pre-trained language model Transformer and the deep learning model DACT as the basis of the malware detection model. This model can reduce the computational load of the large-scale pre-trained language model, balance computational efficiency and computational accuracy, and achieve good results in comprehensive performance. The present invention provides a solution to the problem that the current mainstream malware detection methods are too demanding on device performance and difficult to apply in actual promotion. Description of the Drawings

[0039] Figure 1 It shows the training process and testing process of all experiments provided by the embodiments of the present invention, including data preprocessing, creating a sequence matrix, calculating the output process of each transformer block, etc.;

[0040] Figure 2 It is a model diagram of the Encoder encoder part of the Transformer provided by the embodiments of the present invention;

[0041] Figure 3 It is a model diagram of the internal structure of DACT provided by the embodiments of the present invention;

[0042] Figure 4 It is the model architecture of DACT-Transformer provided by the embodiments of the present invention. Detailed implementation manners

[0043] Combined with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. The embodiments of the present invention provide a malware detection method that combines a large-scale pre-trained language model and deep learning, and the detection process is as follows Figure 1 shown

[0044] First, perform feature screening and feature processing on the original data set, divide the data set according to 7:3, use the training data to train and fine-tune the model, obtain the trained detection model and save it, use the detection model to predict the test set, obtain the prediction results and perform result analysis

[0045] The specific steps are as follows

[0046] Step 1: Data preprocessing

[0047] 1) Delete the rows where the features with most missing values are located, because too many missing values not only have no reference value, but the filled missing values may also have negative impacts

[0048] 2) Find and delete outliers for the same reason as above

[0049] 3) If the missing values and outliers only occupy a small part of a data set, the data set can be filled and modified by the median method

[0050] 4) Feature attribute quantization encoding: Encode the discrete feature attributes within the range of 0-m, where m + 1 represents the total number of software types

[0051] 5) The normalization method adopted is to constrain the feature values using the maximum and minimum values of the current feature, and the formula is as follows

[0052]

[0053] 6) Shuffle the data when dividing the data into training set, validation set and test set to enhance the generalization ability of the model

[0054] Step 2: Use the one-dimensional convolutional neural network Conv1d to expand the two-dimensional data into a multi-dimensional dense vector for the data processed in Step 1, embed the weak association information between the original features into the continuous vector space, and generate a sequence cluster matrix

[0055] Step 3: Put the sequence cluster matrix generated by the convolutional neural network in Step 2 into the model proposed in the present invention for operation. The present invention adopts a malware detection model based on the adaptive computation time strategy - DACT-Transformer

[0056] First, as a large-scale pre-trained language model, Transformer is an architecture that avoids recursion and relies entirely on the attention mechanism to derive the global dependencies between the input and output. Before the Transformer was proposed, most state-of-the-art deep learning models belonged to attention-based Recurrent Neural Network (RNN) models. One drawback of this type of model is that the sequential computational structure makes it unable to process the input sequence in parallel and has a maximum length dependence problem. Another drawback is that the attention mechanism assigns only one importance weight to a feature in the source data, so it can only focus on one aspect of it. Transformer solves these two problems and achieves new state-of-the-art performance relying solely on the improved attention mechanism.

[0057] The Encoder part of Transformer with feature extraction ability is selected in the present invention, such as Figure 2 , which has both the ability to extract long-range dependencies and the ability of parallel computing. These capabilities mainly benefit from the self-attention structure in Transformer-encoder. When calculating, it uses the features before and after it simultaneously so that it can extract the long-range dependencies between features; since the calculation of each feature is independent and does not depend on each other, all features can be calculated in parallel at the same time.

[0058] Secondly, the core of the DACT-Transformer model is the Differentiable Adaptive Computation Time (DACT) algorithm. DACT serves as an additional linear layer as a gating mechanism, and this linear layer can independently learn the confidence function. Then, this mechanism can be used to detect when a stable result is obtained, thereby reducing the total number of steps required for prediction.

[0059] The DACT-Transformer model includes an Embedding layer, N Transformer encoding layers, and N adaptive computation layers DACT; N = 12;

[0060] The Embedding layer is used to expand the input sequence into a feature vector with a dimension of 512 through Embedding;

[0061] The first Transformer encoding layer receives the output of the Embedding layer, and the output after Transformer encoding is copied into two copies, which are respectively input into the next Transformer encoding layer and the DACT layer.

[0062] The n-th Transformer encoding layer receives the output of the previous Transformer encoding layer. The output after Transformer encoding is copied into two parts, which are respectively input into the next Transformer encoding layer and the DACT layer; where 2 ≤ n ≤ N;

[0063] The 1st DACT layer receives the output of the 1st Transformer encoding layer, and after processing the input matrix through Softmax and Sigmoid, it obtains the intermediate prediction result y n and its confidence, and then outputs it to the next DACT layer;

[0064] The n-th DACT layer receives the output of the previous Transformer encoding layer and DACT layer, and after processing the input matrix through Softmax and Sigmoid, it obtains the intermediate prediction result y n and its confidence, and then outputs it to the next DACT layer; In each DACT layer, when the confidence reaches a certain range, the training is terminated, and the current prediction result is output as the intermediate prediction result; 2 ≤ n ≤ N - 1;

[0065] The N-th DACT layer receives the output of the previous Transformer encoding layer and DACT layer, and after processing the input matrix through Softmax and Sigmoid, it obtains the final prediction result y N .

[0066] The Transformer encoding layer includes a MultiHead Attention layer and a fully connected layer. In an encoding layer of a Transformer, first, the input feature vectors are used to calculate the attention distribution through Self-Attention to obtain a matrix. After normalization, it enters a fully connected neural network, and this layer is used to perform a non-linear transformation on the features to improve the expression ability of the model.

[0067] Specifically, the Transformer encoding layer is:

[0068] 1) In a Transformer block, as Figure 2 shown, the sequence cluster matrix generated by the convolutional neural network in step 2 is put into the Multi-head Attention to calculate the attention distribution of the sequence. Among them, the calculation process of the attention mechanism can be roughly divided into 3 steps:

[0069] ① Information input: Receive the input feature vectors q, k, v, where q, k, v are respectively the results of different linear transformations of the matrix X after passing through the embedding layer, and the matrix X = [x 1 , x 2,...,x M is the sequence cluster matrix of step S2 or the output result of the upper-layer Transformer encoding layer; M = 512.

[0070] ② Calculate the attention distribution α: Calculate the relevance through the dot product of q and k, and calculate the score through softmax. Let q = k = v, and calculate the attention weight through softmax. The formula is as shown in (1.2).

[0071] α i = softmax(s(k i , q)) = softmax(s(x i , q)) #(1.11)

[0072] where α i is the attention probability distribution, and s(k i , q) is the attention scoring mechanism.

[0073] ③ Information weighted average: The attention probability distribution α i is used to explain the degree of attention received by the i-th piece of information when querying the context q i . The formula is as shown in (1.3).

[0074]

[0075] where M represents the dimension of the matrix X.

[0076] 4) The sequence matrix calculated through Multi-head Attention reaches the Add&Norm layer. After residual processing and Normalization, it is used to eliminate the information loss problem caused by the deepening of the number of layers.

[0077] 5) The sequence matrix after normalization enters the FeedForward layer. The FeedForward layer is a two-layer fully connected layer. Since in the Multi-head Attention layer, mainly matrix multiplication is performed, which is a linear transformation, and the learning ability of linear transformation is not as good as that of non-linear transformation. Therefore, in the FeedForward layer, the data is first mapped to a high-dimensional space and then mapped to a low-dimensional space, so that more abstract features can be learned.

[0078] 4) The result matrix obtained after another layer of residual processing and normalization process after learning in the FeedForward layer will be copied into two copies, which will be respectively input into the calculation of the sequence matrix of the next layer of Transformer Block, and used to calculate the accuracy of the current model in DACT.

[0079] The specific DACT layer is:

[0080] 1) Inside a DACT block, the intermediate value obtained from the current Transformer block is processed by Softmax and Sigmoid to obtain the intermediate prediction result y n and represents the confidence h of the current module in its output result y n of. h n . h n is the main information providing adaptive computing, as shown in formulas (1.4) and (1.5).

[0081]

[0082]

[0083] where z i represents the i-th classification result among all classification results obtained from the Transformer encoding layer, and C represents the total number of all classification results obtained from the Transformer encoding layer;

[0084] 2) In this way, a value p for assisting in determining the stopping condition can be established n , as shown in (1.6), p n represents the uncertainty considering the entire set of the first n models, and the value of p n must be monotonically decreasing with respect to the index n and tend to 0. When p n reaches a certain range, the training of subsequent model blocks can be stopped. As Figure 3 shown.

[0085]

[0086] P r (ans * , n)(1 - p n ) d ≥ p r (ans ru , n)+p n d#(1.16)

[0087] a n = y n p n-1 + a n-1 (1 - p n-1 )#(1.17)

[0088] The range calculation is as shown in (1.7), and the calculation stops when p n reaches the range. In addition, a nIt is an auxiliary accumulator variable, as shown in (1.8), and is used to combine all intermediate outputs y up to the nth block to form the final result Y. ans * represents the best intermediate output y, d represents the number of remaining steps to distance N, ans ru represents the second-best intermediate output y, P r (ans * , n) represents obtaining ans at the nth layer * probability.

[0089] Step 4: Put the test data into the model trained in Step 3 for detection. The result output in this step is the final prediction result.

[0090] The performance evaluation of the present invention uses the open-source Android malware dataset CICMalDroid 2020, which has more than 17,341 Android samples. The samples in this dataset include five different categories: adware, banking malware, SMS malware, risky software, and benign software.

[0091] The performance evaluation metrics adopted by the present invention are 5 metrics: Accuracy, Precision, Recall, F1-Score, and model training duration (Time / epoch).

[0092] Accuracy refers to the proportion of the number of correctly classified samples to the total number of samples for the given data.

[0093]

[0094] Precision: The proportion of samples that are actually truly positive among the samples predicted as positive.

[0095]

[0096] Recall: The proportion of correctly predicted samples among the samples that are actually truly positive. That is, among the samples that are truly positive, what proportion of the samples are predicted as positive.

[0097]

[0098] The higher the Precision, the stronger the model's ability to distinguish negative samples. The higher the Recall, the stronger the model's ability to distinguish positive samples. F1-Score refers to the harmonic mean of Precision and Recall.

[0099]

[0100] Since Precision and Recall are a pair of conflicting quantities, when P is high, R tends to be relatively low, and when R is high, P tends to be relatively low. Therefore, in order to better evaluate the performance of the classifier, F1-Score is used as the evaluation criterion in the present invention to measure the comprehensive performance of the classifier.

[0101] In addition, this paper also compares the average training duration of each epoch during model training to intuitively measure whether the DACT-Transformer model reduces the computational amount.

[0102] The comparison of the prediction effects of the model of the present invention and other models on the above dataset is shown in Table 1:

[0103] Table 1 Comparison of model prediction effects

[0104]

[0105] It can be seen from Table 1 that in the training task of judging whether it is malware, the Transformer model has the highest accuracy, reaching an average of 97.59%. While the model of the present invention, DACT-Transformer, is only 1.95% lower than the Accuracy of the Transformer model when shortening the average training duration of the model by about 43%; compared with LSTM, when the model Accuracy of the DACT-Transformer model of the present invention is about 1.325% higher, the average training duration of the model is close. Other Precision, Recall, and F1-Score all indicate that the comprehensive performance of the DACT-Transformer model of the present invention is the best.

[0106] The above is the preferred implementation process of the present invention. All changes made according to the technology of the present invention, when the functional effects produced do not exceed the scope of the technical solution of the present invention, fall within the protection scope of the present invention.

Claims

1. A malware detection method based on an adaptive computing time strategy, characterized in that it includes the following steps: S1. Clean the APK data of the Android software with labels, and divide the data to obtain a training set, a validation set, and a test set; the APK data refers to decompressing the APK installation package and parsing the decompression result to extract multiple features of the program, and the multiple features include system calls, binders, and the call frequencies of composite behaviors; where the labels are malware and benign software; S2. Use the convolutional neural network Conv1d to expand the two-dimensional data into a multi-dimensional dense vector for the data processed in step S1, embed the weak association information between the original features into the continuous vector space, and generate a sequence cluster matrix; S3. Build a malware detection model DACT-Transformer based on the adaptive computing time strategy, and use the training set data to train the model; The malware detection model DACT-Transformer based on the adaptive computing time strategy includes an Embedding layer, N Transformer encoding layers, and N adaptive computing layers DACT; where N = 12; The Embedding layer is used to expand the input sequence into a feature vector with a size of 512 dimensions through Embedding; The first Transformer encoding layer receives the output of the Embedding layer, and the output after Transformer encoding is copied into two parts, which are respectively input into the second Transformer encoding layer and the first DACT layer; The nth Transformer encoding layer receives the output of the previous Transformer encoding layer, and the output after Transformer encoding is copied into two parts, which are respectively input into the (n + 1)th Transformer encoding layer and the nth DACT layer; where 2 ≤ n ≤ N - 1; The Nth Transformer encoding layer receives the output of the previous Transformer encoding layer, and the output after Transformer encoding is input into the Nth DACT layer; The first DACT layer receives the output of the first Transformer encoding layer, processes the input matrix through Softmax and Sigmoid to obtain the intermediate prediction result y n and its confidence, and then outputs them to the second DACT layer; The n-th DACT layer is connected to the output of the previous Transformer encoding layer and DACT layer. After processing the input matrix through Softmax and Sigmoid, the intermediate prediction result y n and its confidence are obtained, and then output to the next DACT layer; In each DACT layer, when the confidence level reaches a certain range interval, the training is terminated, and the current prediction result is output as the intermediate prediction result; 2 ≤ n ≤ N - 1; The Nth DACT layer receives the outputs of the Nth Transformer encoding layer and the (N-1)th DACT layer, and after processing the input matrix through Softmax and Sigmoid, obtains the final prediction result y N ; The Transformer encoding layer includes a MultiHead Attention layer and a fully connected layer; the MultiHeadAttention layer is used to calculate the attention distribution of the input feature vector through Self-Attention to obtain a matrix; the fully connected layer is used to perform a non-linear transformation on the features after the output of the MultiHead Attention layer is normalized; S4. Use the trained DACT-Transformer model to detect the test set data; S5. Validate the validation set data using the tested DACT-Transformer model; S6. Utilize the validated DACT-Transformer model to achieve the detection of Android software.

2. The method according to claim 1, characterized in that the method of data cleaning in step S1 is: Clean missing values and outliers, quantitatively encode the feature attributes of the data, and normalize the data.

3. The method according to claim 1, characterized in that the method of data cleaning in step S1 is: 1) Check if there are missing values. If there are, continue to judge whether the number of missing values is greater than the preset threshold A. If so, it is considered that there are too many missing values, and the average or mode of the sample set attributes with the same sample label is selected to fill the missing values. If not, it is considered that there are few missing values, and the data row where the missing value is located can be deleted; 2) Check if there are outliers. If there are, continue to judge whether the number of outliers is greater than the preset threshold B. If so, it is considered that there are too many outliers, and the reasons for the outliers need to be analyzed and the outliers modified. If not, it is considered that there are few outliers, and the data rows including the outliers are deleted; 3) Quantitative encoding of feature attributes: Encode discrete feature attributes within the range 0-m, where m+1 represents the total number of software types; 4) The normalization method adopted is to impose the following constraints on each feature, and the formula is as follows: Among them, X i,norm represents the normalization result of the i-th feature X i , where X i,min and X i,max represent the minimum value and the maximum value of the column where the i-th feature is located, respectively.

4. The method according to claim 1, characterized in that in the malicious software detection model DACT-Transformer based on the adaptive computing time strategy in step S3, the Transformer encoding layer is specifically: 1) The Multi-head Attention layer is used to calculate the attention distribution of the input feature vectors through Self-Attention to obtain a matrix: ① Information input: Receive the input feature vectors q, k, v, where q, k, v are the results after different linear transformations of matrix X respectively, and matrix X = [x 1 , x 2 ,..., x M is the sequence cluster matrix of step S2 or the output result of the previous Transformer encoding layer; ② Calculate the attention distribution α: Calculate the relevance through the dot product of q and k, and calculate the scores through softmax; Let q = k = v, and calculate the attention weights through softmax, and the formula is as shown in (1.2); α i = softmax(s(k i , q)) = softmax(s(x i , q))#(1.2) where α i is the attention probability distribution, and s(k i , q) is the attention scoring mechanism; ③ Information weighted average: attention probability distribution α i to explain the degree of attention received by the i-th piece of information when i the context query is q, as shown in formula (1.3); where M represents the dimension of matrix X; 2) The sequence matrix calculated by the Multi-head Attention reaches the Add&Norm layer, and undergoes residual processing and Normalization standardization; 3) The sequence matrix after standardization processing enters the fully connected layer FeedForward, first maps the data to a high-dimensional space and then maps it to a low-dimensional space to learn more abstract features; 4) The result matrix obtained after passing through a layer of residual processing and standardization process again after the fully connected layer FeedForward ends.

5. The method according to claim 1 or 4, characterized in that in the malicious software detection model DACT-Transformer based on the adaptive computing time strategy in step S3, the DACT layer is specifically: 1) The intermediate prediction result y is obtained after processing the intermediate value obtained from the upper-layer Transformer encoding layer through Softmax and Sigmoid. n And it represents the confidence h of the current layer for its output result y. n n ; where z i represents the i-th classification result among all classification results obtained from the Transformer encoding layer, and C represents the total number of all classification results obtained from the Transformer encoding layer; 2) Establish a value p for assisting in determining the stopping condition n , as shown in (1.6), when p n reaches a certain range, the training of subsequent model blocks can be stopped; P r (ans * ,n)(1 - p n ) d ≥P r (ans ru ,n)+p n d#(1.7) a n = y n p n-1 + a n-1 (1 - p n-1 ) # (1.8) where p n represents the uncertainty of considering the entire set of the first n - 1 layers, and the value of p n is monotonically decreasing with respect to the index n and approaches 0. a n represents an auxiliary accumulator variable used to combine all intermediate outputs y up to the nth DACT layer to form the final result Y, ans * represents the best intermediate output y, d represents the number of remaining steps to N, ans ru represents the second-best intermediate output y, P r (ans * , n) represents the probability of obtaining ans * at the nth layer.

6. A malicious software detection system for implementing the method according to any one of claims 1-5, characterized in that It includes a malware detection model DACT-Transformer based on an adaptive computing time strategy that has been trained, validated, and tested.

7. A computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1-5.

8. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-5 is implemented.