Model training method, malicious file detection method, equipment, medium and program product

By extracting the multimodal features of PDF files and training the classification model, the problem of low accuracy and reliability of the existing malicious PDF file detection scheme is solved, and more efficient malicious file detection is achieved.

CN120030540APending Publication Date: 2025-05-23BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510183190.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing malicious PDF file detection scheme has problems of low accuracy and reliability. It is mainly due to the ineffective sample classification, incomplete feature extraction, and susceptible to noise and outliers.

Method used

By obtaining the binary coded data and content coded data of PDF files, image features and text features are extracted respectively, and the classification model is trained using multimodal feature data to improve the accuracy and reliability of malicious PDF file detection.

Benefits of technology

Through multimodal feature extraction and model training, the accuracy and reliability of malicious PDF file detection are effectively improved, the dependence on a single data source is reduced, and the comprehensiveness and noise resistance of detection are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030540A_ABST
    Figure CN120030540A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method, a malicious file detection method, equipment, a medium and a program product, and relates to the technical field of file detection. The method comprises the following steps: collecting a sample PDF file data set; obtaining binary coded data of the sample PDF file, determining a corresponding transition probability matrix, and converting the matrix into a grayscale image; extracting image features based on the grayscale image; extracting corresponding preprocessing data from the content coding data of the sample PDF file, converting the preprocessing data into word vector data, and extracting text features based on the word vector data; and training a to-be-trained classification model based on the features of the sample PDF file and the corresponding sample labeling information to obtain a trained malicious file detection model. According to the method, the data of different coding forms of the PDF file are obtained, the image features and the text features are extracted from the data, and model training is carried out based on the extracted multi-modal feature data, so that the accuracy and reliability of malicious PDF file detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of file detection technology, and more specifically, to a model training method, a malicious file detection method, a device, a medium, and a program product. Background Art

[0002] In digital communication and document exchange scenarios, PDF files are highly favored due to their good cross-platform compatibility and format retention characteristics. However, the high flexibility of this file format also provides attackers with opportunities to hide malicious code. With the widespread use of PDF files, they have become an important channel for cyber attackers to deliver malware. Attackers can use the structure and functions of PDF files to embed malicious code, including malicious JavaScript code, malicious files, and malicious remote links, to carry out network attacks.

[0003] In actual applications, malicious code detection solutions for PDF files combine static analysis, dynamic analysis, and machine learning. These solutions can achieve intelligent and efficient malicious file detection to a certain extent, but there are still some technical defects, including the use of a single threshold as the basis for detection for linear segmentation, which leads to the inability to effectively classify samples, and incomplete feature extraction, resulting in the loss of key information, making the detection model susceptible to interference from noise and outliers. In summary, the existing malicious PDF file detection solutions have the defects of low accuracy and reliability. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, a malicious file detection method, a device, a medium and a program product to improve the accuracy and reliability of malicious PDF file detection.

[0005] In a first aspect, an embodiment of the present application provides a model training method, comprising:

[0006] Collecting a sample PDF file data set; wherein the sample PDF file data set includes a number of sample PDF files marked as malicious and a number of sample PDF files marked as benign;

[0007] Obtaining binary coded data of each of the sample PDF files, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix;

[0008] Extracting image features corresponding to the sample PDF file based on the grayscale image;

[0009] Obtaining content encoding data of each of the sample PDF files, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule;

[0010] Convert the preprocessed data into word vector data, and extract text features corresponding to the sample PDF file based on the word vector data;

[0011] Based on the image features, text features and corresponding sample annotation information of each of the sample PDF files, the classification model to be trained is trained to obtain a trained malicious file detection model.

[0012] In an embodiment of the present application, by obtaining data in different encoding forms of PDF files and extracting image features and text features therefrom respectively, the extracted multimodal feature data can be used for model training, thereby effectively improving the accuracy and reliability of malicious PDF file detection.

[0013] In some possible embodiments, obtaining binary coded data of each of the sample PDF files, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix includes:

[0014] Obtaining binary coded data of each of the sample PDF files, dividing the binary coded data into states according to a preset number of bits, and counting transition probabilities between states to obtain a transition probability matrix corresponding to the binary coded data;

[0015] Normalizing the transfer probability matrix to obtain a normalized transfer probability matrix;

[0016] The normalized transition probability matrix is ​​converted into a grayscale image.

[0017] In an embodiment of the present application, the accuracy of image feature extraction is further improved by normalizing the transition probability matrix of binary coded data.

[0018] In some possible embodiments, extracting the image features corresponding to the sample PDF file based on the grayscale image includes:

[0019] Based on a pre-built first multi-layer convolutional neural network, feature extraction is performed on the grayscale image to obtain image features corresponding to the sample PDF file; wherein the first multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

[0020] In an embodiment of the present application, a multi-layer convolutional neural network is used to extract features from grayscale images, thereby further improving the accuracy of image feature extraction.

[0021] In some possible embodiments, obtaining the content encoding data of each of the sample PDF files, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule, includes:

[0022] Obtaining content encoding data of each of the sample PDF files, and filtering out first key text information and data to be converted from the content encoding data based on preset key elements;

[0023] The target format conversion method is used to convert the data to be converted, and the data that does not conform to the target format in the converted data is filtered to obtain the second key text information;

[0024] The first key text information and the second key text information are used as preprocessing data corresponding to the content encoding data.

[0025] In the embodiment of the present application, by performing format conversion and data filtering on the selected specific types of data to be converted, the complexity of key text information can be reduced and the comprehensiveness of text feature extraction can be further improved.

[0026] In some possible embodiments, converting the preprocessed data into word vector data includes:

[0027] Constructing word vector training data based on the preprocessed data, and using the word vector training data to train the pre-constructed CBOW model to obtain a trained word vector model;

[0028] The trained word vector model is used to obtain word vector data corresponding to the preprocessed data.

[0029] In an embodiment of the present application, the accuracy of word vector conversion is further improved by constructing training data based on preprocessed data and training a word vector model.

[0030] In some possible embodiments, extracting text features corresponding to the sample PDF file based on the word vector data includes:

[0031] Based on a pre-built second multi-layer convolutional neural network, feature extraction is performed on the word vector data to obtain text features corresponding to the sample PDF file; wherein the second multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

[0032] In an embodiment of the present application, a multi-layer convolutional neural network is used to perform feature extraction on word vector data, thereby further improving the accuracy of text feature extraction.

[0033] In a second aspect, an embodiment of the present application provides a malicious file detection method, including:

[0034] Obtain binary coded data of the PDF file to be detected, determine a transition probability matrix corresponding to the binary coded data in a preset state unit, and convert the data into a grayscale image based on the transition probability matrix;

[0035] Extracting image features corresponding to the PDF file to be detected based on the grayscale image;

[0036] Obtaining content encoding data of the PDF file to be detected, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule;

[0037] Convert the preprocessed data into word vector data, and extract text features corresponding to the PDF file to be detected based on the word vector data;

[0038] The image features and text features of the PDF file to be detected are input into a malicious file detection model trained by any method to obtain a detection result output by the malicious file detection model; wherein the detection result is used to characterize whether the PDF file to be detected is a malicious file.

[0039] In an embodiment of the present application, by obtaining data in different encoding forms of PDF files and extracting image features and text features therefrom respectively, malicious file detection can be performed based on the extracted multimodal feature data, thereby effectively improving the accuracy and reliability of malicious PDF file detection.

[0040] In a third aspect, an embodiment of the present application provides a model training device, comprising:

[0041] A sample collection module, used to collect a sample PDF file data set; wherein the sample PDF file data set includes a number of sample PDF files marked as malicious and a number of sample PDF files marked as benign;

[0042] A sample image conversion module, used for obtaining binary coded data of each sample PDF file, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix;

[0043] A sample image feature extraction module, used to extract image features corresponding to the sample PDF file based on the grayscale image;

[0044] A sample data preprocessing module, used to obtain content coding data of each sample PDF file, and extract preprocessing data corresponding to the content coding data based on a preset processing rule;

[0045] A sample text feature extraction module, used to convert the preprocessed data into word vector data, and extract text features corresponding to the sample PDF file based on the word vector data;

[0046] The model training module is used to train the classification model to be trained based on the image features, text features and corresponding sample annotation information of each sample PDF file to obtain a trained malicious file detection model.

[0047] In a fourth aspect, an embodiment of the present application provides a malicious file detection device, including:

[0048] An image conversion module, used for obtaining binary coded data of a PDF file to be detected, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting the data into a grayscale image based on the transition probability matrix;

[0049] An image feature extraction module, used for extracting image features corresponding to the PDF file to be detected based on the grayscale image;

[0050] A data preprocessing module, used to obtain content coding data of the PDF file to be detected, and extract preprocessing data corresponding to the content coding data based on a preset processing rule;

[0051] A text feature extraction module, used for converting the preprocessed data into word vector data, and extracting text features corresponding to the PDF file to be detected based on the word vector data;

[0052] A malicious file detection module is used to input the image features and text features of the PDF file to be detected into a malicious file detection model trained by any method to obtain a detection result output by the malicious file detection model; wherein the detection result is used to characterize whether the PDF file to be detected is a malicious file.

[0053] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement the method described in any embodiment when executing the program.

[0054] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment can be implemented.

[0055] In a seventh aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, wherein the computer program can implement the method described in any embodiment when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;

[0058] Figure 2 A flowchart of a malicious PDF file detection method provided in an embodiment of the present application;

[0059] Figure 3 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0060] Figure 4 A schematic diagram of the structure of a malicious PDF file detection device provided in an embodiment of the present application;

[0061] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0063] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0064] It should be noted that the current malicious PDF file detection solutions mainly focus on three aspects: static analysis, dynamic analysis and machine learning. Among them, static analysis refers to the detection of malicious code by extracting key features in PDF files, such as metadata, JavaScript code, object structure, etc. without executing the PDF file; or matching features of PDF files based on known malicious features and rule bases to identify malicious files. Dynamic analysis refers to executing PDF files in a controlled environment and monitoring their behavior to detect malicious activities. The application of machine learning in malicious PDF detection is mainly to train classification models by extracting and selecting features that are important for malicious PDF detection, such as byte entropy, number of objects, path features, etc., to achieve automated detection of PDF files.

[0065] Although the current malicious PDF detection technology integrates multiple methods such as static analysis, dynamic analysis and machine learning, it not only improves the security of PDF files, but also provides strong support for network security in multiple industries. However, due to problems such as incomplete feature extraction, the current malicious PDF detection scheme has problems with insufficient accuracy and reliability.

[0066] like Figure 1 As shown, the embodiment of the present application provides a malicious file detection method, which may include the steps of:

[0067] S101, collecting a sample PDF file dataset; wherein the sample PDF file dataset includes a number of sample PDF files marked as malicious and a number of sample PDF files marked as benign;

[0068] Specifically, a certain number of PDF files are collected to construct a data set for training the model, which includes PDF files containing malicious code (malicious PDF files) and normal PDF files, and labels of "malicious" and "benign" are added to the two types of PDF files respectively as annotation information for training.

[0069] S102, obtaining binary coded data of each sample PDF file, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix;

[0070] Specifically, for each sample PDF file in the sample PDF file data set, on the one hand, the binary format data (binary coded data) of the PDF file is obtained, and on the other hand, the programming format data (content coded data) of the PDF file is obtained. Among them, for the binary coded data, firstly, a preset number of bits is used as a state unit, for example, every eight bits (one byte) of binary number is used as a state, and the range of each byte is between 0-255, so each byte has 256 possible states. A transition probability matrix is ​​generated according to the transition probability between states. Exemplarily, the calculation formula of the transition probability matrix is ​​as follows:

[0071]

[0072] Wherein, c[i,j] represents the number of times a byte in state i is followed by a byte in state j, and p[i,j] represents the probability of a byte in state i being transferred to a byte in state j; i, j∈{0,1,2...255}.

[0073] After the transition probability matrix corresponding to the binary coded data is calculated, a corresponding grayscale image can be generated based on the transition probability matrix.

[0074] S103, extracting image features corresponding to the sample PDF file based on the grayscale image;

[0075] Specifically, for the generated grayscale image, a pre-built image feature extraction model may be used to perform feature extraction on the grayscale image to obtain image features corresponding to the sample PDF file.

[0076] S104, obtaining content coding data of each sample PDF file, and extracting pre-processing data corresponding to the content coding data based on a preset processing rule;

[0077] It should be noted that the content encoding data of the sample PDF file refers to the data in the programming format of the PDF file. For example, the encoding data corresponding to the PDF file can be read by a text editor such as UltraEdit. For the obtained content encoding data, corresponding processing rules can be set according to requirements to select and obtain specific information therein as preprocessing data for feature extraction.

[0078] S105, converting the preprocessed data into word vector data, and extracting text features corresponding to the sample PDF file based on the word vector data;

[0079] After the preprocessed data is screened and obtained, it can be converted into a representation of a word vector. For example, the word vector data of the preprocessed data can be obtained by using a bag-of-words model, a word embedding model, etc. For the generated word vector data, a pre-built text feature extraction model can be used to extract features and obtain text features corresponding to the sample PDF file.

[0080] S106: Train the classification model to be trained based on the image features, text features and corresponding sample annotation information of each sample PDF file to obtain a trained malicious file detection model.

[0081] After extracting the feature data (image features and text features) of each sample PDF file, the feature data of the sample PDF file and its corresponding sample annotation information (i.e., the information annotated in step S1) can be used as a piece of training data, and then the collected and processed training data can be divided into a training set and a validation set.

[0082] Among them, the pre-built classification model (exemplarily, an MLP model can be used) is trained iteratively for multiple times using the training set data, and each iteration includes three steps: forward propagation, loss calculation and back propagation. In the forward propagation, the input data is passed through the network to the output layer. Since the input layer receives carefully preprocessed data, these data contain rich semantic information and image texture information; these features are then passed through a series of hidden layers, each of which uses a ReLU activation function to ensure that the network can capture the nonlinear relationship between the features and enhance the generalization ability of the model; the sigmoid activation function can be used in the output layer at the end of the network to output a probability value between 0 and 1, which represents the possibility that the sample belongs to a certain category; in the loss calculation, the loss is calculated based on the difference between the predicted value generated by the output layer and the actual target value; in the back propagation, the gradient descent method can be used to update the weights and biases in the network to minimize the loss function (for example, using cross entropy loss).

[0083] Then, the validation set is used to evaluate the model performance, and the model structure or parameters are adjusted as needed under the drive of minimizing the loss function. Common tuning methods include adjusting the learning rate, increasing or decreasing the number of hidden layers and the number of neurons, etc. Finally, after meeting the preset performance conditions, a trained malicious file detection model is obtained.

[0084] It should be noted that after the image features and text features of the sample PDF file are obtained, the image features and text features may be concatenated to form multimodal fusion features of the sample PDF file.

[0085] Based on this, by obtaining data in different encoding forms of PDF files and extracting image features and text features from them respectively, the model is trained using the extracted multimodal feature data. By complementing the cross-modal feature information of images and text, the dependence on a single data source is effectively reduced, thereby effectively improving the accuracy and reliability of malicious PDF file detection.

[0086] In some possible embodiments, step S102, obtaining binary coded data of each sample PDF file, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix, may include:

[0087] S1021, obtaining binary coded data of each sample PDF file, dividing the binary coded data into states according to a preset number of bits, and counting the transition probabilities between the states to obtain a transition probability matrix corresponding to the binary coded data;

[0088] S1022, normalizing the transfer probability matrix to obtain a normalized transfer probability matrix;

[0089] S1023: Convert the normalized transfer probability matrix into a grayscale image.

[0090] It should be noted that after obtaining the state transition probability p[i, j] corresponding to each byte in the binary coded data, the generated initial probability matrix can be normalized. Exemplarily, the calculation formula for the normalization process is as follows:

[0091]

[0092] Among them, μ and σ represent the mean and standard deviation of all p[i,j] respectively, and p_n[i,j] represents the state transition probability after normalizing p[i,j].

[0093] Then, a grayscale image is generated based on the normalized transition probability matrix, where the pixel value at each position in the image can be expressed as:

[0094] pixel[i,j]=ceil(p_n[i,j]*255)

[0095] Among them, the ceil() function represents a mathematical function that rounds up.

[0096] Based on this, the accuracy of image feature extraction can be further improved by normalizing the transition probability matrix of binary coded data.

[0097] In some possible embodiments, step S103, extracting image features corresponding to the sample PDF file based on the grayscale image, may include:

[0098] S1031. Perform feature extraction on the grayscale image based on a pre-built first multi-layer convolutional neural network to obtain image features corresponding to the sample PDF file; wherein the first multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

[0099] It should be noted that when extracting image features of grayscale images, feature extraction can be achieved by constructing a multi-layer convolutional neural network. For example, a module consisting of N1 [convolutional layers + non-linear activation functions (such as ReLU) + pooling layers] can be used to extract image features of grayscale images. Among them, each convolutional layer is composed of multiple convolutional kernels, and the convolutional kernel performs a convolution operation on the input image by means of a sliding window; after each convolutional layer, a non-linear activation function (such as ReLU) is used to increase the non-linear ability of the neural network; the pooling layer is used to reduce the dimension of the feature map and retain important information.

[0100] Based on this, by using a multi-layer convolutional neural network to extract features from grayscale images, the accuracy of image feature extraction can be further improved.

[0101] In some possible embodiments, step S104, obtaining content encoding data of each sample PDF file, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule, may include:

[0102] S1041, obtaining content encoding data of each sample PDF file, and filtering out first key text information and data to be converted from the content encoding data based on preset key elements;

[0103] S1042, converting the format of the data to be converted using a target format conversion method, and filtering the data that does not conform to the target format in the converted data to obtain second key text information;

[0104] S1043: Use the first key text information and the second key text information as pre-processed data corresponding to the content encoding data.

[0105] It should be noted that by presetting key elements, key information in the content encoding data can be screened, including the first key text information and the data to be converted, wherein the first key text information refers to the data that can be directly used for the next step of text feature extraction, and the data to be converted needs to be format converted first before it can be used for the next step of text feature extraction.

[0106] The preset key element refers to the identifier of the information you want to filter from the content encoding data. For example, the version information of the PDF file can be extracted (usually starting with the element symbol "%"). Secondly, extract the key information starting with a slash ( / ). This information usually represents the keywords in the PDF dictionary and is used to store information. For example, the URI in the PDF can be embedded through the / URI and / Action dictionaries, based on which the complete link information of the obj in the PDF file can be extracted; in addition, JavaScript can be embedded in the PDF through the / JavaScript action, based on which the complete JavaScript content of the obj in the PDF file can be extracted.

[0107] It should be noted that the above key information can be directly used as the first key text information for the next step of text feature extraction. In addition, for the data containing the encoding methods of / ASCIIHexDecode, / ASCII85Decode and / FlateDecode in the Stream, these data are also extracted based on the preset key elements as the data to be converted. For these data, it is necessary to first use the corresponding decoding method (the conversion method of the target format, such as ASCII) to decode them, and filter the data that does not conform to the target format in the decoded data (filter out the characters in non-ASCII format) to obtain the second key text information. In addition, the attacker may also use the complexity of the PDF specification itself to achieve evasion attacks, such as replacing keywords with hexadecimal to cause keyword features to be unable to be extracted, etc. Based on this, these data can also be extracted as the data to be converted based on specific key elements, and the data to be converted is converted based on a specific target format, such as converting a hexadecimal sequence into ASCII text, etc.

[0108] For the extracted first key text information and second key text information, these data information may be concatenated using an empty character string as preprocessing data corresponding to the content encoding data.

[0109] Based on this, by performing format conversion and data filtering on the selected specific types of data to be converted, the complexity of key text information can be reduced and the comprehensiveness of text feature extraction can be further improved.

[0110] In some possible embodiments, in step S105, converting the preprocessed data into word vector data may include:

[0111] S1051, constructing word vector training data based on the preprocessed data, and using the word vector training data to train the pre-constructed CBOW model to obtain a trained word vector model;

[0112] S1052. Use the trained word vector model to obtain word vector data corresponding to the preprocessed data.

[0113] It should be noted that the word embedding model based on the CBOW architecture can be used to convert the preprocessed data into word vectors. Specifically, the preprocessed data is first processed by word segmentation, punctuation removal, conversion to lowercase, and stop words removal, and then these data are randomly initialized into word vector form as word vector training data; then, the CBOW model is constructed, and the parameters such as the dimension of the word vector, the size of the context window, and the minimum word frequency of the sampling size are initialized; during the training process, for each training sample, a central word is selected, and then n surrounding words are selected from the context of the central word as input, where n is determined by the size of the context window; then the average value of the selected context word vectors or concatenation of them is used to predict the central word through a neural network, which usually contains one or more hidden layers; then a loss function (usually a cross entropy loss function) is used to measure the difference between the predicted central word and the actual central word; finally, the word vector is adjusted through the back propagation algorithm to minimize the loss function, and this process will update the vector representation of each word in the word embedding table.

[0114] After the CBOW model is trained, the trained word vectors can be obtained from the embedding layer of the model to obtain the word vector data corresponding to the preprocessed data.

[0115] Based on this, by constructing training data based on the preprocessed data and training the word vector model to obtain the word vector data corresponding to the preprocessed data, the accuracy of the word vector conversion is further improved.

[0116] In some possible embodiments, in step S105, extracting text features corresponding to the sample PDF file based on word vector data may include:

[0117] S1053. Perform feature extraction on the word vector data based on a pre-built second multi-layer convolutional neural network to obtain text features corresponding to the sample PDF file; wherein the second multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

[0118] It should be noted that when extracting text features of word vector data, feature extraction can be achieved by constructing a multi-layer convolutional neural network. For example, similar to the feature extraction of grayscale images, a module consisting of N2 [convolutional layers + non-linear activation functions (such as ReLU) + pooling layers] can be used to extract text features of word vector data. Among them, each convolutional layer is composed of multiple convolutional kernels, and the convolutional kernel performs convolution operations on the primary text features input by this layer in the form of a sliding window; after each convolutional layer, a non-linear activation function (such as ReLU) is used to increase the non-linear ability of the neural network; the pooling layer is used to reduce the dimension of the feature data and retain important information.

[0119] Based on this, a multi-layer convolutional neural network is used to extract features from word vector data, thereby further improving the accuracy of text feature extraction.

[0120] like Figure 2 As shown, the embodiment of the present application provides a malicious file detection method, which may include:

[0121] S201, obtaining binary coded data of a PDF file to be detected, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting the data into a grayscale image based on the transition probability matrix;

[0122] S202, extracting image features corresponding to the PDF file to be detected based on the grayscale image;

[0123] S203, obtaining content coding data of the PDF file to be detected, and extracting preprocessing data corresponding to the content coding data based on a preset processing rule;

[0124] S204, converting the preprocessed data into word vector data, and extracting text features corresponding to the PDF file to be detected based on the word vector data;

[0125] S205. Input the image features and text features of the PDF file to be detected into a malicious file detection model trained by any one of the methods to obtain a detection result output by the malicious file detection model; wherein the detection result is used to indicate whether the PDF file to be detected is a malicious file.

[0126] It should be noted that in the process of executing the malicious file detection method, for the PDF file to be detected, Figure 1 The model training method shown is similar, and the same data processing and feature extraction process is required. Finally, the extracted image and text cross-module feature information is input into the trained malicious file detection model. The model will output the detection result indicating whether the PDF file to be detected is a malicious file.

[0127] It can be understood that the above-mentioned malicious file detection method corresponds to the embodiment of the model training method of the present invention. For the convenience and conciseness of description, the specific working process of the malicious file detection method described above can refer to the corresponding process in the aforementioned model training method, and will not be elaborated here.

[0128] Please refer to Figure 3 , Figure 3 FIG. 1 shows a block diagram of a model training device provided in some embodiments of the present application. It should be understood that the model training device is similar to the above-mentioned Figure 1 Corresponding to the method embodiment, it is able to execute each step involved in the above method embodiment. The specific functions of the model training device can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0129] Figure 3 The model training device includes at least one software function module that can be stored in a memory or fixed in the model training device in the form of software or firmware, and the model training device includes:

[0130] The sample collection module 310 is used to collect a sample PDF file data set; wherein the sample PDF file data set includes a number of sample PDF files marked as malicious and a number of sample PDF files marked as benign;

[0131] The sample image conversion module 320 is used to obtain binary coded data of each sample PDF file, determine a transition probability matrix corresponding to the binary coded data in a preset state unit, and convert it into a grayscale image based on the transition probability matrix;

[0132] A sample image feature extraction module 330 is used to extract image features corresponding to the sample PDF file based on the grayscale image;

[0133] The sample data preprocessing module 340 is used to obtain the content coding data of each sample PDF file and extract the preprocessing data corresponding to the content coding data based on a preset processing rule;

[0134] A sample text feature extraction module 350 is used to convert the preprocessed data into word vector data, and extract text features corresponding to the sample PDF file based on the word vector data;

[0135] The model training module 360 ​​is used to train the classification model to be trained based on the image features, text features and corresponding sample annotation information of each sample PDF file to obtain a trained malicious file detection model.

[0136] It can be understood that the above-mentioned device item embodiments correspond to the method item embodiments of the present invention. A model training device provided by the embodiments of the present invention can implement the model training method provided by any method item embodiment of the present invention.

[0137] Please refer to Figure 4 , Figure 4 FIG. 1 is a block diagram showing a malicious PDF file detection device provided by some embodiments of the present application. It should be understood that the malicious PDF file detection device is similar to the above-mentioned Figure 2 Corresponding to the method embodiment, each step involved in the above method embodiment can be executed. The specific functions of the malicious PDF file detection device can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0138] Figure 4 The malicious PDF file detection device includes at least one software function module that can be stored in a memory in the form of software or firmware or fixed in the malicious PDF file detection device, and the malicious PDF file detection device includes:

[0139] An image conversion module 410 is used to obtain binary coded data of a PDF file to be detected, determine a transition probability matrix corresponding to the binary coded data in a preset state unit, and convert the data into a grayscale image based on the transition probability matrix;

[0140] An image feature extraction module 420 is used to extract image features corresponding to the PDF file to be detected based on the grayscale image;

[0141] The data preprocessing module 430 is used to obtain the content coding data of the PDF file to be detected, and extract the preprocessing data corresponding to the content coding data based on the preset processing rules;

[0142] A text feature extraction module 440 is used to convert the preprocessed data into word vector data, and extract text features corresponding to the PDF file to be detected based on the word vector data;

[0143] The malicious file detection module 450 is used to input the image features and text features of the PDF file to be detected into the malicious file detection model trained by any method to obtain the detection result output by the malicious file detection model; wherein the detection result is used to characterize whether the PDF file to be detected is a malicious file.

[0144] It can be understood that the above-mentioned device item embodiment corresponds to the method item embodiment of the present invention. A malicious PDF file detection device provided by the embodiment of the present invention can implement the malicious file detection method provided by any method item embodiment of the present invention.

[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0146] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 reads the program from the memory 510 through a bus 530 and executes the program to implement a method of any embodiment included in the above-mentioned malicious file detection method.

[0147] Processor 520 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 520 can be a microprocessor.

[0148] The memory 510 may be used to store instructions executed by the processor 520 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present application. The processor 520 of the disclosed embodiment may be used to execute instructions in the memory 510 to implement the method shown above. The memory 510 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.

[0149] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon. The computer program is executed by a processor to execute the method described in the method embodiment.

[0150] Some embodiments of the present application further provide a computer program product, which, when executed on a computer, enables the computer to execute the method described in the method embodiment.

[0151] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0152] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0153] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0154] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0155] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0156] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0157] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A model training method, characterized in that: include: Collecting a sample PDF file data set; wherein the sample PDF file data set includes a number of sample PDF files marked as malicious and a number of sample PDF files marked as benign; Obtaining binary coded data of each of the sample PDF files, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting into a grayscale image based on the transition probability matrix; Extracting image features corresponding to the sample PDF file based on the grayscale image; Obtaining content encoding data of each of the sample PDF files, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule; Convert the preprocessed data into word vector data, and extract text features corresponding to the sample PDF file based on the word vector data; Based on the image features, text features and corresponding sample annotation information of each of the sample PDF files, the classification model to be trained is trained to obtain a trained malicious file detection model.

2. The model training method according to claim 1, characterized in that: The step of obtaining binary coded data of each of the sample PDF files, determining a transition probability matrix corresponding to the binary coded data in a preset state unit, and converting the transition probability matrix into a grayscale image includes: Obtaining binary coded data of each of the sample PDF files, dividing the binary coded data into states according to a preset number of bits, and counting transition probabilities between states to obtain a transition probability matrix corresponding to the binary coded data; Normalizing the transfer probability matrix to obtain a normalized transfer probability matrix; The normalized transition probability matrix is ​​converted into a grayscale image.

3. The model training method according to claim 1, characterized in that: The extracting the image features corresponding to the sample PDF file based on the grayscale image includes: Based on a pre-built first multi-layer convolutional neural network, feature extraction is performed on the grayscale image to obtain image features corresponding to the sample PDF file; wherein the first multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

4. The model training method according to claim 1, characterized in that: The obtaining of the content coding data of each of the sample PDF files and extracting the pre-processing data corresponding to the content coding data based on a preset processing rule includes: Obtaining content encoding data of each of the sample PDF files, and filtering out first key text information and data to be converted from the content encoding data based on preset key elements; The target format conversion method is used to convert the data to be converted, and the data that does not conform to the target format in the converted data is filtered to obtain the second key text information; The first key text information and the second key text information are used as preprocessing data corresponding to the content encoding data.

5. The model training method according to claim 1, characterized in that: The converting the preprocessed data into word vector data comprises: Constructing word vector training data based on the preprocessed data, and using the word vector training data to train the pre-constructed CBOW model to obtain a trained word vector model; The trained word vector model is used to obtain word vector data corresponding to the preprocessed data.

6. The model training method according to claim 1, characterized in that: The extracting text features corresponding to the sample PDF file based on the word vector data includes: Based on a pre-built second multi-layer convolutional neural network, feature extraction is performed on the word vector data to obtain text features corresponding to the sample PDF file; wherein the second multi-layer convolutional neural network includes a plurality of modules consisting of convolutional layers, non-linear activation functions and pooling layers.

7. A malicious file detection method, characterized in that: include: Obtain binary coded data of the PDF file to be detected, determine a transition probability matrix corresponding to the binary coded data in a preset state unit, and convert the data into a grayscale image based on the transition probability matrix; Extracting image features corresponding to the PDF file to be detected based on the grayscale image; Obtaining content encoding data of the PDF file to be detected, and extracting preprocessing data corresponding to the content encoding data based on a preset processing rule; Convert the preprocessed data into word vector data, and extract text features corresponding to the PDF file to be detected based on the word vector data; The image features and text features of the PDF file to be detected are input into a malicious file detection model trained by the method according to any one of claims 1 to 6 to obtain a detection result output by the malicious file detection model; wherein the detection result is used to characterize whether the PDF file to be detected is a malicious file.

8. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement any method of claims 1-7 when executing the program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Sample cleaning method for AI or ML model

    CN120256970A