Coded file intrusion detection method and system based on CAE-XGBoost

By adopting the CAE-XGBoost method in malicious file detection, using a convolutional autoencoder to extract depth features and combining SMOTE and XGBoost, the problem of insufficient detection capabilities for unknown malicious files in the existing technology is solved, and it is more adaptable and robust, which is suitable for the detection needs of complex data sets.

CN120145382APending Publication Date: 2025-06-13BEIJING KEDONG ELECTRIC POWER CONTROL SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346384.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing malicious file detection methods show significant limitations when facing unknown malicious files and mutant viruses, which are difficult to effectively identify and detect, and are highly dependent on manual feature engineering.

Method used

The encoded file intrusion detection method based on CAE-XGBoost is adopted, and the depth features are automatically extracted from the encoded files through a convolutional autoencoder, and combined with the SMOTE oversampling algorithm and the XGBoost integrated learning module, efficient detection of unknown malicious files is achieved.

Benefits of technology

This method can automatically learn deep features from complex coded files, reduce dependence on manual feature selection, improve the difference in feature distribution of different types of coded files, enhance the adaptability and robustness of detection, and can efficiently process large-scale data sets, which is suitable for coding file detection scenarios across safe zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145382A_ABST
    Figure CN120145382A_ABST
Patent Text Reader

Abstract

The invention discloses a coding file intrusion detection method and system based on CAE-XGBoost. The method comprises the following steps: in a data preparation stage, collecting coding files from a plurality of data sources and carrying out standardization processing on the coding files to serve as input data of an intrusion detection model; training a CAE model by using the input data until an encoder of the CAE model is solidified, and performing depth feature extraction on the input data by using the CAE model with the solidified encoder; in the classification training stage, using an SMOTE oversampling algorithm to increase a minority class of samples, and then using an XGBoost integrated learning module to classify the extracted depth features to obtain a classification result of the depth features; and based on a classification result, evaluating the performance of the intrusion detection model by using an accuracy rate, a macro average accuracy rate and a recall rate index. According to the scheme of the invention, the adaptability and robustness of intrusion detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of file detection, and particularly relates to an encoding file intrusion detection method and system based on CAE-XGBoost. Background Art

[0002] Traditional malicious file detection methods mainly rely on rules and signature matching. These methods build a feature library by pre-collecting the characteristics of known malicious files, such as file hash values, specific byte sequences, or code fragments, and compare and match the files to be detected to identify known threats. Such methods are simple to implement and have a fast detection speed. However, this rule-based detection method has significant limitations when facing unknown malicious files and variant viruses. Rule detection often fails to effectively identify them, the detection effect drops significantly, and it is easily bypassed by attackers.

[0003] To make up for the deficiencies of rule-based methods, machine learning-based malicious file detection technologies have gradually emerged. These methods extract features from files, such as byte entropy, opcode sequences, API call frequencies, etc., and use machine learning algorithms to train classification models to achieve the detection of unknown malicious files. Commonly used machine learning algorithms include Support Vector Machine (SVM), Random Forest, Naive Bayes, etc. Traditional machine learning algorithms have limited performance when dealing with high-dimensional, non-linear, and complex data, and rely heavily on manual feature engineering. If the feature selection is inappropriate, it will greatly affect the model effect and lack generalization. Summary of the Invention

[0004] In order to solve the deficiencies existing in the prior art, the present invention provides an encoding file intrusion detection method and system based on CAE-XGBoost to solve the technical problem of coping with security threats in the scenario of encoding file transmission.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions.

[0006] The present invention first discloses an encoding file intrusion detection method based on CAE-XGBoost, and the method includes the following steps:

[0007] Step 1: In the data preparation stage, collect encoding files from multiple data sources and perform normalization processing as the input data of the intrusion detection model;

[0008] Step 2: Use the input data to train the CAE model until the encoder of the CAE model is solidified, and use the CAE model with the solidified encoder to perform deep feature extraction on the input data;

[0009] Step 3: In the classification training stage, use the SMOTE oversampling algorithm to increase the minority class samples, and then use the XGBoost ensemble learning module to classify the extracted deep features to obtain the classification results of the deep features;

[0010] Step 4: Based on the classification results, use accuracy, macro-average precision, and recall metrics to evaluate the performance of the intrusion detection model.

[0011] The present invention further includes the following preferred solutions:

[0012] The collecting and normalizing the encoded files from multiple data sources further includes:

[0013] Collect various types of encoded files from various data sources. The encoded files include executable files, various scripts, dynamic link libraries, and white data for training and testing the intrusion detection model;

[0014] Convert the collected files through visible character encoding;

[0015] Normalize the encoded files, including duplicate removal, character index encoding, and length normalization, so that all data has a consistent structure and format;

[0016] Divide the processed input data into a training set and a test set.

[0017] The training the CAE model using the input data until the encoder of the CAE model is solidified, and using the CAE model with the solidified encoder to extract deep features from the input data further includes:

[0018] Perform one-hot conversion on the encoded files after integer index conversion;

[0019] Use the normal encoded file sample data in the training set data to train the CAE model, and solidify the CAE encoder part after reaching the preset training effect;

[0020] Input all data sets into the CAE encoder to obtain the latent space encoded feature data as the input of the subsequent classification model, and add corresponding type labels to the training set. The type labels include binary executable files, scripts, and white samples.

[0021] The using the SMOTE oversampling algorithm to increase the minority class samples, and then using the XGBoost ensemble learning module to classify the extracted deep features further includes:

[0022] Perform data oversampling on the three types of minority classes in the training set, and use the SMOTE oversampling method to generate new minority class samples;

[0023] Perform cross-validation on the training data and optimize the hyperparameters of the model through grid search;

[0024] Use the optimized hyperparameter combination to train the XGBoost model, and persistently store the trained XGBoost model for the classification task of the test set.

[0025] The classification algorithm is implemented by a random forest or gradient boosting tree classification algorithm.

[0026] The encoder of the CAE model consists of a convolutional layer, a pooling layer, and a Dropout layer. Through layer-by-layer convolution and pooling operations, it extracts and compresses the feature representation of the input data to generate a compact feature suitable for the decoder to process; the convolutional layer is used to extract the local patterns of the input data, the pooling layer is used to reduce the dimension and enhance the robustness of the model, and the Dropout layer prevents the model from overfitting.

[0027] The performing cross-validation on the training data and optimizing the hyperparameters of the model through grid search further includes:

[0028] Use three-fold cross-validation to obtain the best hyperparameter model:

[0029] 1) max_depth: The maximum depth of a single model tree is 7;

[0030] 2) min_child_weight: The minimum node weight is 3;

[0031] 3) gamma: 0.3, used to control overfitting;

[0032] 4) lambda: 2, used to control overfitting;

[0033] 5) Number of base models: 500;

[0034] 6) Specify the optimization objective objective: multi softmax to select multi-class output.

[0035] The present invention also discloses a CAE-XGBoost-based encoded file intrusion detection system using the foregoing CAE-XGBoost-based encoded file intrusion detection method, including:

[0036] A data preparation module, used in the data preparation stage to collect encoded files from multiple data sources and perform normalization processing as the input data of the intrusion detection model;

[0037] A convolutional auto-encoding module, used to train the CAE model using the input data until the encoder of the CAE model is solidified, and use the CAE model with the solidified encoder to perform deep feature extraction on the input data;

[0038] A classification training module, which is used to increase the minority class samples using the SMOTE oversampling algorithm during the classification training phase, and then use the XGBoost ensemble learning module to classify the extracted deep features to obtain the classification results of the deep features;

[0039] A performance evaluation module, which is used to evaluate the performance of the intrusion detection model based on the classification results using accuracy, macro-average precision, and recall metrics.

[0040] Correspondingly, the present application also discloses a terminal, including a processor and a storage medium;

[0041] The storage medium is used to store instructions;

[0042] The processor is used to operate according to the instructions to execute the steps of the foregoing CAE-XGBoost-based encoded file intrusion detection method.

[0043] Correspondingly, the present application also discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the foregoing CAE-XGBoost-based encoded file intrusion detection method.

[0044] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention provides a CAE-XGBoost-based encoded file intrusion detection method and system. The convolutional autoencoder can automatically learn deep features from complex encoded files, enhancing the feature distribution differences of different types of encoded files, especially suitable for data with high complexity and non-linear features, and reducing the dependence on manual feature selection; XGBoost enables very fast training and detection speeds through multi-threaded parallel processing and a block-based storage structure, and can operate efficiently even on large-scale data sets. For the encoded file detection scenario across security zones, where the data volume is large and the real-time detection requirement is high, the detection model using parallel processing is more suitable for the scenario requirements; by modularly combining feature extraction, data augmentation, and classification models, a complete intrusion detection system is formed. This system not only has stronger adaptability and robustness, but also can cope with diverse network attack means and improve the overall protection effect. Description of the Drawings

[0045] Figure 1 is a flowchart of the CAE-XGBoost-based encoded file intrusion detection method in the present invention.

[0046] Figure 2 is a schematic diagram of the overall implementation of the intrusion detection model in the present invention.

[0047] Figure 3It is a schematic diagram of the encoder network structure of 1D-CAE in the present invention. Specific implementation manners

[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0049] The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Aiming at the deficiencies of the prior art, the intrusion detection system in this article is applied to the network security protection scenario of transmission services in different security zones of the power system. It is required that binary executable files, scripts and other high-risk files are not allowed to be transmitted from a low-security zone to a high-security zone, and only current business data, pictures, videos, compressed packages, etc. are allowed to pass through. In addition, the transmission between different security zones requires encrypted communication through a private protocol, and visible character encoding is performed when encrypting and packaging the transmitted files.

[0051] First, in a normal power industrial control environment, collect the encoded files of various normal services, use these encoded files to train the CAE model, unsupervised learn the hidden data patterns from the normal encoded files, and solidify the encoder model of the CAE. Then, construct encoded files of different high-risk file types, and use the CAE encoder model to automatically extract deep data features from all training data, effectively enhancing the feature distribution differences of different categories of encoded files. Next, use the Synthetic Minority Over-sampling Technique (SMOTE) for data augmentation to generate a new training data set. Finally, use the XGBoost model for classification, and through its powerful ensemble learning ability and efficient parallel computing performance, achieve efficient and real-time detection of hidden malicious encoded files.

[0052] The encoding file intrusion detection method based on CAE-XGBoost disclosed in the present invention includes the following steps:

[0053] Step 1: In the data preparation stage, collect encoded files from multiple data sources and perform normalization processing as the input data of the intrusion detection model.

[0054] The intrusion detection model for encoded files based on CAE-XGBoost consists of three parts: the CAE convolutional autoencoder module, the SMOTE oversampling module, and the XGBoost integrated machine learning module. The convolutional autoencoder module obtains the intermediate encoded features of multi-modal encoded files; the SMOTE oversampling module is used to balance different types of data in the training data; the XGBoost integrated learning module is used to perform classification training on the encoded features and obtain a classification detection model. The specific structural form is as Figure 1 shown. Figure 2 The overall implementation process of the intrusion detection method based on CAE-XGBoost is shown.

[0055] The data preparation process includes the following steps:

[0056] Step 1.1: Collect various types of encoded files from various data sources, including high-risk files such as executable files, various scripts, dynamic link libraries, etc., as well as white data such as pictures, videos, compressed packages, etc., for the training and testing of the intrusion detection model.

[0057] Step 1.2: Convert the collected files through visible character encoding to simulate the common evasion strategies of attackers. Encoding methods such as Base16, Base32, and Base64 are used.

[0058] Step 1.3: Normalize the encoded files, including deduplication, character index encoding, and length normalization, etc., to ensure that all data has a consistent structure and format.

[0059] Step 1.4: Divide the processed input data into a training set and a test set to ensure independence during model training and evaluation.

[0060] Step 2: Use the input data to train the CAE model until the encoder of the CAE model is solidified, and use the CAE model with the solidified encoder to perform deep feature extraction on the input data.

[0061] Deep feature extraction is the core part of the entire intrusion detection method. The structure of the CAE model contains two parts: an encoder and a decoder. The encoder is used to convert high-dimensional input data into low-dimensional latent feature vectors, and the decoder attempts to reconstruct the original input so that the encoder can learn effective feature representations. The encoder part that can output hidden layer features is finally solidified in the model, and the implementation process is as follows:

[0062] Step 2.1: Encode the input data of the CAE model: Perform one-hot conversion on the encoded file after character index conversion. Since the index value of each feature does not represent size but represents a visible character type, in order to let the model understand the input data, it needs to be extended to the corresponding dimension.

[0063] Step 2.2: Solidify the encoder of the CAE model: Use the normal encoding file sample data in the training data to train the CAE model, and solidify the CAE encoder part after achieving the preset training effect;

[0064] CAE consists of two parts, an encoder and a decoder, which are used for feature encoding and data reconstruction respectively. Only the detailed formula of its encoder principle is described below:

[0065] The goal of the encoder is to compress the input data X ∈ R n into a low-dimensional feature vector Z ∈ R m (where m < n). Its main process is to perform convolution operations on the input signal through multiple one-dimensional convolutional layers

[15] . Each convolution operation can be expressed as:

[0066]

[0067] where x j+i-1 represents the j + i - 1-th value of the input signal; w i is the weight of the convolution kernel, with a length of k; b is the bias term; f(.) is the activation function, usually the ReLU function, expressed as:

[0068] f(x) = max(0, x)

[0069] Through one-dimensional convolution operations, the encoder can extract the local patterns of the input data and gradually reduce the dimension of the input, and obtain a compact representation z through training.

[0070] Step 2.3: Feature extraction output: Input all data sets into the CAE encoder to obtain the latent space encoded feature data, which is used as the input of the subsequent classification model, and add corresponding type labels to the training set. In a specific embodiment, the type labels include binary executable files, scripts, and white samples.

[0071] Step 3: In the classification training stage, use the SMOTE oversampling algorithm to increase the minority class samples, and then use the XGBoost ensemble learning module to classify the extracted deep features to obtain the classification results of the deep features.

[0072] The specific process of Step 3 further includes:

[0073] Step 3.1: Perform data oversampling on the three types of minority classes (usually malicious samples) in the training set, and use the SMOTE oversampling method to generate new minority class samples to increase the number of samples to solve the problem of data imbalance. In an alternative embodiment, data oversampling uses other data augmentation methods such as random oversampling.

[0074] Specifically, for each minority-class sample x i , SMOTE will randomly select k of its neighbors x i1 x i2 ...x ik , and generate new samples through interpolation. SMOTE can increase the number of minority-class samples to make it more balanced with the majority-class samples.

[0075] For each selected neighbor x j (j = 1, 2,..., k), a new sample x i is generated between the sample x j and its neighbor x new according to the following formula:

[0076] x new = x i + λ · (x j - x i )

[0077] λ is a random value that satisfies 0 ≤ λ ≤ 10 and is used to control the position of the new sample. This value is randomly generated from a uniform distribution to ensure that the new sample is located between x i and x j .

[0078] Step 3.2: Optimize the hyperparameters of the model through the Grid Search and cross-validation methods, systematically search different combinations of hyperparameters, and determine key parameters such as the best tree depth, learning rate, regularization coefficient, etc. Specifically, use 10-fold cross-validation to evaluate the performance of each set of hyperparameters, and select the combination of hyperparameters with the optimal performance on the validation set to ensure that the model can achieve the best generalization ability and classification accuracy in the classification task.

[0079] Step 3.3: Train the XGBoost model using the optimized combination of hyperparameters, and persistently store the trained XGBoost model for the classification task of the test set.

[0080] XGBoost is an improved gradient boosting decision tree (GBDT) algorithm, and its core idea is to improve the model performance by constructing multiple weak decision trees. Specifically, it continuously iteratively adds new trees to fit the residuals of the previous iteration to optimize the objective function.

[0081] Its objective function consists of a loss function and a regularization term, and the formula is:

[0082]

[0083] Among them, the loss function is used to measure the prediction error, and the regularization term Control the model complexity to prevent overfitting. In the formula, T is the number of leaf nodes of the k-th tree; w j is the weight of the j-th leaf node; γλ is a hyperparameter used to adjust regularization to prevent overfitting.

[0084] In each round of iteration, the model is updated by minimizing the negative gradient. The gain measures the quality of the split, and the split point with the largest gain is selected. The gain formula is:

[0085]

[0086] In the formula, G L and G R represent the sum of gradients of the left and right child nodes respectively; H L and H R represent the sum of second-order derivatives of the left and right child nodes respectively;

[0087] In an alternative embodiment, the classification model can also adopt models such as random forest and gradient boosting tree.

[0088] Step 4: Based on the classification results, use accuracy, macro-average precision, and recall metrics to evaluate the performance of the intrusion detection model.

[0089] After the classification training is completed, the performance of the model is evaluated. Since it is a multi-classification task, the three metrics of accuracy, macro-average precision, and recall are selected. Through their combination, the detection ability of the intrusion detection model and its performance in practical applications are comprehensively measured to ensure the effective detection of malicious files, while reducing false positives and false negatives, thereby optimizing the overall network security protection ability.

[0090] In a specific embodiment of the present invention, a multi-layer 1D CNN is used to implement the compressed representation of the encoded file sequence samples, and the convolutional kernel size is set to 3. Figure 3 The network structure of the encoder part is shown in detail in

[0091] Input layer: The shape of the input data is 64×3000×256, where 64 represents the batch size, 3000 represents the feature sequence length, and 256 represents the input feature dimension, that is, the integer index encoding dimension.

[0092] Convolutional layer 1: The first convolutional layer receives the input data and maps it to 64×3000×64. One-dimensional convolution operation is performed through 64 convolutional kernels to extract local features from the input data.

[0093] Pooling Layer 1: Next, the first pooling layer performs a pooling operation to reduce the dimensionality of the convolutional output to 64×1500×64. The pooling layer reduces the dimensionality of the feature map by taking the maximum or average value of a local window, thereby reducing the computational amount and enhancing the translational invariance of the features.

[0094] Dropout Layer: The pooled data passes through the Dropout layer, which is used to randomly discard a part of the neurons to prevent overfitting of the model and maintain the generalization ability of the features. The output shape of this layer is still 64×1500×64.

[0095] Convolutional Layer 2: After passing through Dropout, the data is passed to the second convolutional layer, and the output is 64×1500×32, that is, the number of channels is reduced from 64 to 32 to further compress the feature representation and extract higher-level features.

[0096] Pooling Layer 2: Finally, through the second pooling layer, the dimensionality is reduced to 64×750×32, further compressing the length of the feature data for subsequent encoding processes.

[0097] This encoder structure effectively extracts and compresses the feature representation of the input data through successive convolutional and pooling operations, generating compact features suitable for further processing by the decoder. The convolutional layer is used to extract local patterns of the input data, while the pooling layer is used to reduce the dimensionality and enhance the robustness of the model, and Dropout prevents overfitting of the model and improves the generalization ability.

[0098] The decoder structure of the CAE is exactly the mirror image of the encoder, increasing the convolutional kernel layer by layer, and the Dropout layer becomes an oversampling layer. Through these operations, the CAE can efficiently encode the input data while retaining its main features and potential information. The specific complete model structure of the CAE is shown in Table 1:

[0099] Table 1

[0100]

[0101] Regarding the pattern features of the encoded file data obtained through the CAE network structure, use Grid Search to find the best combination of model hyperparameters on the training set, and use three-fold cross-validation to obtain the following best hyperparameter model:

[0102] 1) max_depth: The maximum depth of the single-model tree is 7. Among the encoded features obtained by deep feature learning of the input, the single-model leaf node can be split up to 7 times at most, which can prevent the model from overfitting the training set and resulting in insufficient generalization ability of the model;

[0103] 2) min_child_weight: 3, the minimum node weight is used to control the generation of leaf nodes;

[0104] 3) gamma: 0.3, which is used to control overfitting;

[0105] 4) lambda: 2, which is used to control overfitting. Since overfitting was found during debugging, a relatively large value is set here;

[0106] 5) Number of base models: 500. Using 500 trees for comprehensive determination can increase the overall generalization ability of the model;

[0107] 6) Specify the optimization objective objective: multi softmax to select multi-classification output.

[0108] In an alternative embodiment, the CAE convolutional autoencoder can be replaced by a similar deep learning model, such as an LSTM autoencoder, a denoising autoencoder, etc.

[0109] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention provides a CAE-XGBoost-based coding file intrusion detection method and system. The convolutional autoencoder can automatically learn deep features from complex coding files, enhancing the feature distribution differences of different types of coding files, especially suitable for data with high complexity and non-linear features, and reducing the dependence on manual feature selection; XGBoost enables very fast training and detection speeds through multi-threaded parallel processing and a block-based storage structure, and can operate efficiently even on large-scale data sets. For the coding file detection scenario across security zones, where the data volume is large and the real-time detection requirement is high, the detection model using parallel processing is more suitable for the scenario requirements; by modularly combining feature extraction, data augmentation, and classification models, a complete intrusion detection system is formed. This system not only has stronger adaptability and robustness, but also can cope with diverse network attack means, improving the overall protection effect.

[0110] The present invention can be a system, a method, and / or a computer program product. The present invention also discloses a CAE-XGBoost-based coding file intrusion detection system based on the aforementioned CAE-XGBoost-based coding file intrusion detection method, including:

[0111] A data preparation module, which is used to collect coding files from multiple data sources and perform normalization processing during the data preparation stage as the input data of the intrusion detection model;

[0112] A convolutional autoencoding module, which is used to train the CAE model using the input data until the encoder of the CAE model is solidified, and use the CAE model with the solidified encoder to perform deep feature extraction on the input data;

[0113] A classification training module, which is used in the classification training stage to increase the minority class samples using the SMOTE oversampling algorithm, and then use the XGBoost ensemble learning module to classify the extracted deep features to obtain the classification results of the deep features;

[0114] A performance evaluation module, which is used to evaluate the performance of the intrusion detection model based on the classification results using accuracy, macro-average precision, and recall metrics.

[0115] Based on the spirit of the present invention, those skilled in the art can easily conceive that a computer program product can be obtained based on the foregoing CAE-XGBoost-based encoded file intrusion detection method. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure. That is, the present application also includes a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the foregoing CAE-XGBoost-based encoded file intrusion detection method.

[0116] A computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example, - but not limited to - an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding devices, such as punch cards or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.

[0117] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0118] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A coding file intrusion detection method based on CAE-XGBoost, characterized in that: The following steps are involved: Step 1: In the data preparation phase, the encoded files are collected from multiple data sources and normalized as input data for the intrusion detection model; Step 2: training the CAE model using the input data until the encoder of the CAE model is solidified, and performing deep feature extraction on the input data using the CAE model with the encoder solidified; Step 3: In the classification training stage, the SMOTE oversampling algorithm is used to increase the minority class samples, and then the XGBoost integrated learning module is used to classify the extracted deep features to obtain the classification results of the deep features; Step 4: Based on the classification results, the performance of the intrusion detection model is evaluated using accuracy, macro-average precision, and recall metrics.

2. The CAE-XGBoost-based encoding file intrusion detection method according to claim 1, characterized in that: The collecting of encoded files from multiple data sources and performing normalization processing further includes: Collect various types of encoded files from various data sources, including executable files, various scripts, dynamic link libraries, and white data for training and testing intrusion detection models; Convert the collected files into visible character encoding format; Normalize the encoded files, including deduplication, character index encoding, and length normalization, so that all data has a consistent structure and format; The processed input data is divided into training set and test set.

3. The CAE-XGBoost-based encoding file intrusion detection method according to claim 2 is characterized in that: The step of training the CAE model using the input data until the encoder of the CAE model is solidified, and extracting deep features from the input data using the CAE model whose encoder is solidified, further includes: Perform one-hot conversion on the encoded file after integer index conversion; Use the normal encoding file sample data in the training set data to train the CAE model, and solidify the CAE encoder part after achieving the preset training effect; All data sets are input into the CAE encoder to obtain latent space encoding feature data as the input of the subsequent classification model, and corresponding type labels are added to the training set respectively. The type labels include binary executable files, scripts and white samples.

4. The CAE-XGBoost-based encoding file intrusion detection method according to claim 3 is characterized in that: The method of using the SMOTE oversampling algorithm to increase minority class samples and then using the XGBoost ensemble learning module to classify the extracted deep features further includes: Oversample the data of the three types of minority classes in the training set and use the SMOTE oversampling method to generate new minority class samples; Cross-validate the training data and optimize the model’s hyperparameters through grid search; Use the optimized hyperparameter combination to train the XGBoost model, and store the trained XGBoost model persistently for the classification task of the test set.

5. The CAE-XGBoost-based encoding file intrusion detection method according to claim 4 is characterized in that: The classification algorithm is implemented by a random forest or gradient boosting tree classification algorithm.

6. The CAE-XGBoost-based encoding file intrusion detection method according to claim 5, characterized in that: The encoder of the CAE model consists of a convolutional layer, a pooling layer and a Dropout layer. Through layer-by-layer convolution and pooling operations, the feature representation of the input data is extracted and compressed to generate compact features suitable for decoder processing; the convolutional layer is used to extract the local pattern of the input data, the pooling layer is used to reduce the dimension and enhance the robustness of the model, and the Dropout layer prevents the model from overfitting.

7. The CAE-XGBoost-based encoding file intrusion detection method according to claim 6 is characterized in that: The cross-validation of the training data and optimization of the hyperparameters of the model by grid search further include: Use three-fold cross validation to get the best hyperparameter model: 1) max_depth: The maximum depth of a single model tree is 7; 2)min_child_weight: the minimum node weight is 3; 3)gamma: 0.3, used to control overfitting; 4) lambda: 2, used to control overfitting; 5) Number of base models: 500; 6) Specify the optimization objective:multi softmax and select multi-classification output.

8. A CAE-XGBoost-based encoded file intrusion detection system, characterized in that: include: The data preparation module is used to collect and normalize the coded files from multiple data sources during the data preparation phase as input data for the intrusion detection model; A convolutional autoencoder module, used to train the CAE model using the input data until the encoder of the CAE model is solidified, and to perform deep feature extraction on the input data using the CAE model with the encoder solidified; The classification training module is used to increase the minority class samples using the SMOTE oversampling algorithm during the classification training phase, and then classify the extracted deep features using the XGBoost ensemble learning module to obtain the classification results of the deep features; The performance evaluation module is used to evaluate the performance of the intrusion detection model based on the classification results using accuracy, macro-average precision and recall indicators.

9. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the CAE-XGBoost-based encoding file intrusion detection method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the CAE-XGBoost-based encoding file intrusion detection method described in any one of claims 1 to 7 are implemented.