Malicious file detection method, device, equipment and storage medium

By performing image processing and text recognition of the image information of the to be processed, and combining the two types of recognition results, we can determine whether there is macro deception behavior in the file, which solves the shortcomings of the existing detection methods in terms of universality and accuracy, and achieves more efficient detection of malicious macro deception files.

CN114021137BActive Publication Date: 2025-05-23QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111424941.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-05-23
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

The existing malicious macro file detection methods have shortcomings in generality and detection accuracy. Attackers can bypass detection by modifying the variable name or code obfuscation.

Method used

By performing image processing and text recognition of the image information of the file to be processed, and combining two types of recognition results, we can determine whether there is macro deception in the file. The specific steps include obtaining the image information of the file, performing image processing and text recognition, and determining whether the file is a malicious inducing file based on the processing results and text characteristics.

Benefits of technology

It improves the detection accuracy of malicious macros to trick files, reduces the probability of attackers bypassing detection, and enhances the universality of detection methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021137B_ABST
    Figure CN114021137B_ABST
Patent Text Reader

Abstract

The present application provides a malicious inducement file detection method, device, equipment and storage medium, the method comprising: obtaining image information of a file to be processed; performing image processing on the image information to obtain an image processing result of the image information, and performing text recognition on the image information to obtain text features in the image information; determining whether the file to be processed is a malicious inducement file based on the image processing result and the text features. The present application utilizes the fundamental characteristics of macro deception files, classifies image processing results and recognizes text content on the preview image of the file homepage at the same time, and combines the two types of recognition results to determine whether the file has macro deception behavior, thereby improving the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to a method, device, equipment and storage medium for detecting malicious files. Background Art

[0002] Since the advent of the Internet era, various cybercrimes have occurred frequently, and malicious file attacks, as one of the common means of cybercrime, have also received more and more attention. As a widely used file format worldwide, Office files are the first to be affected and have become an important tool for criminals to package malicious operations. Another study shows that 98% of malicious office files mainly use macros to perform malicious operations. Macro operations are a series of automated operations performed through macro commands. Office (referring to the file format used by Microsoft's Office series of office software) files can store some macro commands (which can be understood as some automated programs), which can call system resources to perform a series of operations. The original intention of macro operations is to improve the efficiency of file editing, but because of the system permissions it has, it has become an important way for criminals to carry out malicious attacks and is often used to perform some malicious operations.

[0003] In order to defend against malicious macro operations, Office 2007 and later versions have disabled the default macro-enabling setting, that is, only when the user actively clicks to enable macros, can macro commands be executed. This allows most macro-deception files to use various means to induce users to enable macros. For example, when a file is opened in Office software, the software detects that the file contains macros, so a security warning will appear at the top of the file page: "Macros are disabled." In order to make users believe that the file is harmless, attackers often use some official icons and lies such as "to protect file information" and "to display the file normally" to induce users to click the "Enable content" button. Once the user is deceived and clicks, the macro code will start to run automatically to perform some malicious operations.

[0004] Existing malicious macro file detection methods are mainly based on static code analysis. The most classic method is to extract the macro code of the file for sensitive word matching. Some studies have also combined the knowledge of natural language processing to analyze and extract the macro code to find the difference in the semantic features of malicious and non-malicious macro code. Since these methods analyze the code, it is easy for attackers to circumvent them by modifying variable names, obfuscating the code, etc., and they are often designed only for certain data sets or file types, and are relatively poor in versatility. Summary of the invention

[0005] The purpose of the embodiments of the present application is to provide a malicious file detection method, device, equipment and storage medium, which utilizes the fundamental characteristics of macro-deception files to simultaneously classify image processing results and recognize text content on the file homepage preview image, and combine the two types of recognition results to determine whether the file contains macro deception behavior, thereby improving detection accuracy.

[0006] A first aspect of an embodiment of the present application provides a malicious-induced file detection method, comprising: obtaining image information of a file to be processed; performing image processing on the image information to obtain an image processing result of the image information, and performing text recognition on the image information to obtain text features in the image information; and determining whether the file to be processed is a malicious-induced file based on the image processing result and the text features.

[0007] In one embodiment, performing image processing on the image information to obtain the image processing result of the image information includes: inputting the image information into a preset recognition model, and outputting the image processing result of the image information, wherein the preset recognition model is at least used to extract image features from the image information.

[0008] In one embodiment, the method further includes: obtaining a sample file data set, selecting a training set and a test set from the sample file data set, wherein each sample file in the training set and the test set is labeled with a label indicating whether it is a malicious inducement file; using the training set to train a neural network model, and using the test set to test the trained model, to obtain the preset recognition model.

[0009] In one embodiment, the training set is used to train the neural network model, and the test set is used to test the trained model to obtain the preset recognition model, including: using the training set to train the neural network model to obtain a primary classification model; using the test set to test the primary classification model, collecting a set of error sample sets of the primary classification model on the test set, wherein the recognition results of the samples in the error sample set are different from the labels of the corresponding samples in the test set; selecting similar samples whose similarity with the error sample reaches a first threshold from the remaining sample file data set, and the remaining sample file data set is the data set after the training set is removed from the sample file data set; adding the similar samples to the training set, and using the updated training set to train the neural network model, iteratively updating the training set until the preset recognition model whose test results reach a preset accuracy rate is established.

[0010] In one embodiment, performing text recognition on the image information to obtain text features in the image information includes: performing text recognition on the image information to obtain text content in the image information as the text features in the image information; and / or performing text recognition on the image information to obtain text content in the image information; extracting word vectors of the text content, and based on the word vectors, extracting semantic features of the text content to obtain text features in the image information.

[0011] In one embodiment, the image processing result includes a first probability that the file to be processed is a malicious inducement file; the text feature includes the text content in the image information; determining whether the file to be processed is a malicious inducement file based on the image processing result and the text feature includes: judging whether there is an identification word for malicious inducement files in the text content based on the text feature; if the identification word exists in the text content, determining that the file to be processed is a malicious inducement file; if the identification word does not exist in the text content, judging whether the first probability is greater than or equal to a preset probability threshold; if the first probability is greater than or equal to the preset probability threshold, determining that the file to be processed is a malicious inducement file, otherwise, determining that the file to be processed is not a malicious inducement file.

[0012] In one embodiment, the image processing result includes image features of the file to be processed; determining whether the file to be processed is a malicious inducement file based on the image processing result and the text features includes: fusing the image features and the text features to generate a fusion feature of the file to be processed; determining a second probability that the file to be processed is a malicious inducement file based on the fusion feature; if the second probability is greater than or equal to a preset probability threshold, determining that the file to be processed is a malicious inducement file, otherwise, determining that the file to be processed is not a malicious inducement file.

[0013] The second aspect of an embodiment of the present application provides a malicious-induced file detection device, including: a first acquisition module, used to acquire image information of a file to be processed; an identification module, used to perform image processing on the image information to obtain an image processing result of the image information, and perform text recognition on the image information to obtain text features in the image information; a determination module, used to determine whether the file to be processed is a malicious-induced file based on the image processing result and the text features.

[0014] In one embodiment, the recognition module is used to: input the image information into a preset recognition model, and output the image processing result of the image information, and the preset recognition model is at least used to extract image features from the image information.

[0015] In one embodiment, it also includes: a second acquisition module, used to obtain a sample file data set, select a training set and a test set from the sample file data set, each sample file in the training set and the test set is marked with a label whether it is a malicious inducement file; an establishment module, used to use the training set to train a neural network model, and use the test set to test the trained model to obtain the preset recognition model.

[0016] In one embodiment, the establishment module is used to: use the training set to train the neural network model to obtain a primary classification model; use the test set to test the primary classification model, collect the wrong sample set of the primary classification model on the test set, the recognition results of the samples in the wrong sample set are different from the labels of the corresponding samples in the test set; select similar samples whose similarity with the wrong sample reaches a first threshold from the remaining sample file data set, and the remaining sample file data set is the data set after the training set is removed from the sample file data set; add the similar samples to the training set, and use the updated training set to train the neural network model, iteratively update the training set until the preset recognition model with the test result reaching the preset accuracy is established.

[0017] In one embodiment, the recognition module is also used to: perform text recognition on the image information to obtain text content in the image information as the text feature in the image information; and / or perform text recognition on the image information to obtain text content in the image information; extract word vectors of the text content, and based on the word vectors, extract semantic features of the text content to obtain text features in the image information.

[0018] In one embodiment, the image processing result includes a first probability that the file to be processed is a malicious-induced file; the text feature includes the text content in the image information; the determination module is used to: determine whether there is an identification word for malicious-induced files in the text content based on the text feature; if the identification word exists in the text content, determine that the file to be processed is a malicious-induced file; if the identification word does not exist in the text content, determine whether the first probability is greater than or equal to a preset probability threshold; if the first probability is greater than or equal to the preset probability threshold, determine that the file to be processed is a malicious-induced file, otherwise, determine that the file to be processed is not a malicious-induced file.

[0019] In one embodiment, the image processing result includes the image features of the file to be processed; the determination module is used to: fuse the image processing result and the text features to generate the fused features of the file to be processed; determine the second probability that the file to be processed is a malicious inducement file based on the fused features; if the second probability is greater than or equal to a preset probability threshold, determine that the file to be processed is a malicious inducement file, otherwise, determine that the file to be processed is not a malicious inducement file.

[0020] A third aspect of an embodiment of the present application provides an electronic device, comprising: a memory for storing a computer program; a processor for executing the computer program to implement the method of the first aspect of the embodiment of the present application and any embodiment thereof.

[0021] The fourth aspect of the embodiments of the present application provides a non-transitory electronic device readable storage medium, including: a program, which, when executed by an electronic device, enables the electronic device to execute the method of the first aspect of the embodiments of the present application and any one of its embodiments.

[0022] The malicious file detection method, device, equipment and storage medium provided by the present application utilize the fundamental characteristics of macro decoy files to simultaneously classify image processing results and recognize text features of the image information of the file to be processed, and combine the two recognition results to determine whether the file contains macro decoy behavior. Compared with the prior art, it not only reduces the probability of attackers circumventing detection by changing the code and improves the versatility of the detection method, but also combines the image processing results and the text feature recognition method to improve the detection accuracy of macro-induced attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A schematic diagram of the structure of an electronic device according to an embodiment of the present application;

[0025] Figure 2A This is a schematic diagram of an example of a macro deception file according to an embodiment of the present application;

[0026] Figure 2B This is a schematic diagram of an example of a macro attack but without textual meaning in one embodiment of the present application;

[0027] Figure 3A schematic diagram of a flow chart of a malicious file detection method according to an embodiment of the present application;

[0028] Figure 4A A schematic diagram of a flow chart of a malicious file detection method according to an embodiment of the present application;

[0029] Figure 4B A schematic diagram of a single training process of an image classification model according to an embodiment of the present application;

[0030] Figure 4C A schematic diagram of a similar sample search process according to an embodiment of the present application;

[0031] Figure 4D A schematic diagram of the detailed structure of the MobileNetV3 model of an embodiment of the present application;

[0032] Figure 4E This is a schematic diagram of a feature vector extraction process according to an embodiment of the present application;

[0033] Figure 5 A schematic diagram of a flow chart of a malicious file detection method according to an embodiment of the present application;

[0034] Figure 6 A schematic diagram of the structure of a malicious file detection device according to an embodiment of the present application. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the present application, the terms "first", "second" and the like are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0036] like Figure 1 As shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12, Figure 1 A processor is taken as an example. The processor 11 and the memory 12 are connected via a bus 10. The memory 12 stores instructions that can be executed by the processor 11, and the instructions are executed by the processor 11 so that the electronic device 1 can execute all or part of the process of the method in the following embodiment to improve the detection accuracy of malicious inducement files and the versatility of the detection method.

[0037] In one embodiment, the electronic device 1 may be a mobile phone, a tablet computer, a laptop computer, a desktop computer, or a large computing system composed of multiple computers.

[0038] In order to more clearly describe the technical content of this embodiment, the application scenarios involved are exemplified as follows:

[0039] Office file: refers to the file format used by Microsoft's Office series of office software, mainly including Compound File Binary Format (CFB) and Office Open XML Format (OOXML). CFB is the standard file format used by Office 97 to 2003, including doc / xls / ppt and other files, which can be decomposed into multiple independent data stream files. OOXML is the mainstream structure used in Office2007 and later, including docx, xlsx, pptx, etc. It can be regarded as an unzipped ZIP package. Users can unzip and obtain the pictures, macro codes and other contents inside without opening the file.

[0040] In order to prevent malicious macro operations, Office 2007 and later versions have disabled the default macro-enabling setting, that is, macro commands can only be executed if the user actively clicks to enable macros. This means that most macro-deceptive files have to use pictures or text to guide users to enable macros. Figure 2A The figure below shows a malicious file opened in Office software. The software detects that the file contains macros, so a security warning will appear at the top of the file page: "Macros are disabled". In order to convince users that the file is harmless, attackers often use official icons and lie such as "to protect file information" or "to make the file display normally" to trick users into clicking the "Enable content" button. Once users are deceived and click the button, the macro code will automatically run to perform some malicious operations.

[0041] In response to the above problems, the known research is basically based on code analysis and identification, and there are few studies that detect malicious files based on images. For example, there is a paper on malicious file detection based on images, "Analysis and Correlation of Visual Evidence in Campaigns of Malicious Office Documents". The technical content described in this paper still has the following defects:

[0042] The above-mentioned paper decompresses the Office file, obtains the image contained in the file, and then performs text recognition on the image. When the sample uses a pure text method to induce, the image cannot be decompressed from the file, and the method mentioned in the paper will fail. In addition, some attack methods are to put a picture with macro attack but no text meaning to the Office series of office software. Meaningless pictures such as blurred pictures, or pictures with only garbled characters, or pictures with only symbols such as arrows, as shown in Figure 2B, the file only contains meaningless garbled content. In this case, the software automatically pops up "Enable content" because it detects the existence of macros. The user may think that "I did not enable the content, which caused the file to display garbled / hidden", and thus fall into the trap and click to start the macro command. Therefore, it only relies on text recognition of pictures, which is very easy to miss.

[0043] Please see Figure 3 , which is a malicious inducement file detection method according to an embodiment of the present application, the method can be Figure 1 The electronic device 1 shown in the figure can be used to perform Figure 2A In the malicious file detection scenario shown in the figure, the detection accuracy of malicious files and the versatility of the detection method are improved. The method can be used on a server or a client. This embodiment takes the server as an example. The method includes the following steps:

[0044] Step 301: Obtain image information of a file to be processed.

[0045] In this step, the file to be processed may be the Office file mentioned in the above scenario, and the image information of the file to be processed may be the preview image of the home page of the Office file, which can be obtained through a preview image acquisition tool, such as an Oracle tool. In actual scenarios, not all Office files can be decompressed to obtain images. For example, the above-mentioned referenced paper method can only detect Office files in OOXML format, because only files in OOXML format can be decompressed to obtain images. However, this embodiment is not limited to the ability of the file to be processed to decompress images, and the home page preview image can also be applied, so it is applicable to all types of Office files, expanding the scope of application.

[0046] Step 302: performing image processing on the image information to obtain an image processing result of the image information, and performing text recognition on the image information to obtain text features in the image information.

[0047] In this step, the image information can be input into a preset recognition model, and the image processing result of the image information can be output. The preset recognition model is an image classification model based on deep learning. The preset recognition model is at least used to extract image features from the image information. After the image information of the to-be-processed file is input into the preset recognition model, the image processing result of the image information can be obtained. Here, the image processing result may include shallow features and deep features of the image information. At the same time, the image information of the to-be-processed file can be subjected to text recognition through the OCR recognition model, and text features can be obtained. The text features may at least include the text position and text content in the image information.

[0048] Step 303: Determine whether the file to be processed is a malicious misleading file based on the image processing result and text features.

[0049] In this step, simply relying on the image processing results or text features of the file to be processed may result in missed detections. Therefore, comprehensively considering the information dimensions represented by the image processing results and text features can more comprehensively display the true attributes of the file to be processed. According to the image processing results and text features, it is determined whether the file to be processed is a malicious or misleading file. For malicious or misleading files that are difficult to identify from the image processing results, text features can be used to ensure the overall detection capability. Similarly, for malicious or misleading files that are difficult to identify from the text feature level, image processing results can also be used to ensure the overall detection capability.

[0050] The above-mentioned malicious induced file detection method utilizes the fundamental characteristics of macro decoy files, simultaneously classifies the image processing results and recognizes the text features of the image information of the processed file, and combines the two recognition results to determine whether the file contains macro decoy behavior. Compared with the existing technology, it not only reduces the probability of attackers circumventing detection by changing the code and improves the versatility of the detection method, but also combines the image processing results and the text feature judgment method to improve the detection accuracy of macro-induced attacks.

[0051] Please see Figure 4A , which is a malicious inducement file detection method according to an embodiment of the present application, the method can be Figure 1 The electronic device 1 shown in the figure can be used to perform Figure 2A In the malicious file detection scenario shown in the figure, the detection accuracy of malicious files and the versatility of the detection method are improved. The method includes the following steps:

[0052] Step 401: Obtain image information of the file to be processed. For details, refer to the description of step 301 in the above embodiment.

[0053] Step 402: Input the image information into a preset recognition model and output the image processing result of the image information. For details, refer to the description of step 302 in the above embodiment.

[0054] In one embodiment, before step 402, a step of establishing the above-mentioned preset recognition model may also be included, as follows:

[0055] S1: Obtain a sample file dataset, select a training set and a test set from the sample file dataset, and each sample file in the training set and the test set is labeled with whether it is a malicious inducement file.

[0056] In this embodiment, the initial data set may include malicious files captured by the protection software and file data collected online. In the data set processing flow, samples with macros can be first screened out from the initial data set, and then the screened samples are deduplicated using md5. Assuming that the image information is the first page preview image of the file, the first page preview image set of each sample file is obtained through the preview image acquisition tool. Finally, the first page preview image set is deduplicated using md5 to obtain the final sample file data set, denoted as U. N pictures are randomly selected from the data set U as the training set A and annotated. Each sample file in the training set is labeled with a label indicating whether it is a malicious induced file. K pictures are then extracted from the remaining sample file data set UA excluding the training set A as the test set T1. Each sample file in the test set T1 is labeled with a label indicating whether it is a malicious induced file.

[0057] S2: Use the training set to train the neural network model, and use the test set to test the trained model to obtain a preset recognition model.

[0058] In this step, a lot of achievements have been made in the fields of image classification and text recognition based on deep learning, such as the neural network structures of VGG series, ResNet series, MobileNet series, etc. Considering the speed and performance, the classic lightweight network MobileNetV3 can be used in this embodiment to extract and classify image features. The model has a small amount of calculation, fast inference speed, and is easy to deploy on various platforms. The neural network model MobileNetV3 is trained with a training set, and the trained model is tested with a test set to obtain a preset recognition model.

[0059] In one embodiment, step S2 may specifically include: using a training set to train a neural network model to obtain a primary classification model. Using a test set to test the primary classification model, collecting a set of false positive samples of the primary classification model on the test set, wherein the recognition results of the samples in the false positive sample set are different from the labels of the corresponding samples in the test set. Selecting similar samples whose similarity with the false positive samples reaches a first threshold from the remaining sample file data set, wherein the remaining sample file data set is a data set after the training set is removed from the sample file data set. Adding similar samples to the training set, and using the updated training set to train the neural network model, iteratively updating the training set until a preset recognition model is established whose test results reach a preset accuracy rate.

[0060] In the actual model training process, the data set can be expanded based on similar sample search. Figure 4B As shown in Figure 1, we can first use the training set A to train the MobileNetV3 network structure to obtain the primary classification model M1, and then extract K pictures from the remaining sample file dataset UA excluding the training set A as the test set T1. For the wrong sample set [ws 1 ,ws 2 …ws n ] (where n is a positive integer), and through the similar sample search method, search ws in the remaining sample file dataset UA in sequence i (i is an integer greater than 1 and less than or equal to n) similar samples, and add the similar samples to the training set A, and repeat the above process n times until the model test result of model Mn on the training set Tn reaches the preset accuracy.

[0061] In one embodiment, for the search process of similar samples, such as Figure 4C As shown in Figure 1, for each sample image in the sample file dataset U (assuming a images), the model M can be used to calculate its feature vector, and the feature vectors of all images are integrated and saved as a feature set D (size a×576). i When analyzing, first extract the wrong example sample ws i The feature vector (size is 1×576), and then the wrong sample ws i The similarity is calculated between the feature vector of and all the feature vectors in the feature set D. That is, for each sample image us in the dataset U i , you can get a us i With ws i Finally, we take all the similarity scores s greater than or equal to the first threshold us i As ws isimilar samples, wherein the first threshold is a similarity threshold, which can be set based on actual needs, for example, the first threshold can be 0.6.

[0062] In one embodiment, if the sample image us has a similarity score s greater than or equal to 0.6 i There are many samples, which can be sorted from high to low based on similarity, and the sample images with higher rankings can be selected. i As ws i For example, we select sample images with a similarity greater than 0.6 and ranked in the top 20. i As ws i The above method of expanding the data set with similar samples, when the data annotation cost is high, can first annotate a small number of samples (such as training set A), and train the primary model based on a small number of samples, and then search for similar samples on a large amount of unlabeled data according to the image features of the test error samples, and select more meaningful samples to add to the training set, so as to quickly improve the model effect.

[0063] In one embodiment, if Figure 4D As shown, the structure of the MobileNetV3-small network model may include a feature extraction module and a feature classification module, wherein the feature extraction module is used to extract image features from image information, and the feature classification module is used to determine the probability that the file to be processed is a malicious inducement file based on the image features. Each module includes a specific convolutional layer structure. In this embodiment, the output of the last pooling layer in the feature extraction module of the MobileNetV3-small network model can be used as the feature vector of each image, and the specific dimension can be 1x576.

[0064] like Figure 4E As shown in FIG. 1 , it is a schematic diagram of the feature vector extraction process. After the preset recognition model is established, the image information of the file to be processed is first input into the feature extraction module of the MobileNetV3-small network model, and the feature vector of the input image is obtained from the output result of the feature extraction module. On the other hand, the output result of the feature extraction module is continuously input into the feature classification module, and the classification result of the input image is output. That is, the image processing result may include the first probability that the file to be processed belongs to a malicious decoy file, for example, the first probability black_p that the file to be processed belongs to a macro decoy file (recorded as a black sample) and the probability white_p that the file to be processed belongs to a non-macro decoy file (recorded as a white sample) may be output, where the sum of black_p and white_p is 1. Of course, only the probability that the file to be processed belongs to a macro decoy file may be output as needed, or only the probability that the file to be processed belongs to a non-macro decoy file may be output.

[0065] Step 403: Perform text recognition on the image information to obtain the text content in the image information as the text feature in the image information.

[0066] In this step, the text features include the text content in the image information, and the text position detection and content recognition can be performed on the first page preview image of the processed file through a commonly used OCR recognition model to obtain the text content in the form of a string.

[0067] It should be noted that the image processing process in step 402 and the text feature recognition in step 403 can be performed simultaneously or sequentially, and this embodiment does not limit the order of their implementation.

[0068] Step 404: Determine whether there is a malicious file identification word in the text content based on the text features. If yes, proceed to step 406; otherwise, proceed to step 405.

[0069] In this step, common malicious inducement files will carry special identification words, such as "Enable Macro", "Enable Content" and other identification words with similar meanings. The identification words of malicious inducement files can be counted in advance to obtain a keyword library, and then the text recognition results of step 403 are matched with keywords. If the identification words in the keyword library are hit, go to step 406, otherwise go to step 405.

[0070] Step 405: Determine whether the first probability is greater than or equal to a preset probability threshold. If yes, proceed to step 406; otherwise, proceed to step 407.

[0071] In this step, if the text content does not contain the identification words in the keyword library, in order to avoid missing some malicious deceptive files that use non-text, such as blurred picture files, the text content cannot be detected, but its image characteristics are very deceptive. Therefore, the classification result of the preset recognition model in step 402 is further judged, that is, whether the first probability black_p of the to-be-processed file belonging to the macro deceptive file is greater than or equal to the preset probability threshold, if so, go to step 406, otherwise go to step 407. Among them, the preset probability threshold can be determined based on the actual situation, for example, it can be 0.98.

[0072] Step 406: Determine whether the file to be processed is a malicious file.

[0073] In this step, if there are identification words in the text content of the file to be processed, or although there are no identification words in the keyword library in the text content, but the first probability black_p of the file to be processed being a macro deception file in the image feature classification result of the file to be processed is greater than or equal to the preset probability threshold, then the file to be processed can be directly determined to be a malicious inducement file.

[0074] Step 407: Determine whether the file to be processed is a malicious file.

[0075] In this step, if the text content of the file to be processed does not contain any identification words in the keyword library, and the first probability black_p of the file to be processed being a macro decoy file in its image feature classification result is less than a preset probability threshold, it can be determined that the file to be processed is not a malicious decoy file.

[0076] The above malicious decoy file detection method simultaneously classifies image features and identifies text content on the homepage preview of a file containing macro code, and combines the two types of results to determine whether the file contains macro decoy behavior. The types of files to be processed referred to in this method include but are not limited to Office files. In fact, all files that meet the characteristics of malicious decoy files can be applied to this method. Its advantages are as follows:

[0077] (1) Different from the traditional method that mainly analyzes based on easily modifiable codes, the above embodiment directly analyzes the contents of the file based on the basic characteristics of the macro decoy file. It is very difficult for an attacker to bypass this detection method without reducing the success rate of the attack. Precisely because macro decoy files must use misleading content to lure users to open macros, the method of this embodiment can, to a certain extent, achieve "unchanging in the face of changes", and the update cycle required for the model is much longer than that of traditional methods. Practice has proved that during the online deployment monitoring phase of more than one month, the detection rate of malicious samples has always remained at a high level, and the overall false positive rate is low.

[0078] (2) In this embodiment, a lightweight deep learning model is used, which is small in size, fast in speed, easy to deploy, and has low hardware requirements.

[0079] (3) The method in this embodiment has nothing to do with the Office file format. It only needs to extract the file preview image. It can be applied to Office files in any format and is very versatile.

[0080] Please see Figure 5 , which is a malicious inducement file detection method according to an embodiment of the present application, the method can be Figure 1 The electronic device 1 shown in the figure can be used to perform Figure 2A In the malicious file detection scenario shown in the figure, the detection accuracy of malicious files and the versatility of the detection method are improved. The method includes the following steps:

[0081] Step 501: Obtain image information of the file to be processed. For details, refer to the description of step 301 in the above embodiment.

[0082] Step 502: Input the image information into a preset recognition model, and output an image processing result of the image information, wherein the image processing result includes the image features of the file to be processed. For details, refer to the description of step 302 and the model building step in step 402 in the above embodiment.

[0083] Step 503: Perform text recognition on the image information to obtain the text content in the image information. For details, see the description of step 403 in the above embodiment.

[0084] Step 504: extract word vectors of the text content, and based on the word vectors, extract semantic features of the text content to obtain text features in the image information.

[0085] In this step, the text features include the text content and semantic features in the image information. Natural language processing related technologies can be used to extract word vectors from the text content, and then semantic features are obtained based on the word vectors. Semantic features are the mapping of text in a high-dimensional feature space, similar to image features to images. Compared with direct keyword matching, this semantic-based analysis method will make the discrimination results more accurate and more robust.

[0086] Step 505: fusing the image features and the text features to generate fused features of the file to be processed.

[0087] In this step, the image processing result includes the image features of the file to be processed, and the semantic features in the text features are fused with the image features obtained in step 502 to generate fused features of the file to be processed. That is, the preset recognition model includes a feature fusion module at this time, and feature fusion can adopt early fusion: first fuse multiple layers of features, and then train the classifier on the fused features. Two classic feature fusion methods that can be used are as follows:

[0088] (1) concat: Serial feature fusion, directly concatenating two features. If the dimensions of the two input features x and y are p and q, the dimension of the output feature z is p+q.

[0089] (2) add: Parallel strategy, combining the two feature vectors into a complex vector. For input features x and y, z = x + iy, where i is the imaginary unit.

[0090] Step 506: Determine a second probability that the file to be processed is a malicious inducement file based on the fused features.

[0091] In this step, a layer of neural network classifier can be added after the feature extraction module and feature fusion module of the preset recognition model. The structure of the classifier can be, for example, a fully connected layer + softmax function. By inputting the fused features into the classifier, the second probability that the file to be processed is a malicious inducement file can be obtained. The classifier calculates the second probability that the file to be processed is a malicious inducement file based on the received fused features, and then classifies the file to be processed as a malicious inducement file or a non-malicious inducement file based on the second probability.

[0092] In this embodiment, the preset recognition model includes at least: a feature extraction module, a feature fusion module and a classifier. The model training method can be similar to FIG. 4A to FIG. 4E The method shown in will not be repeated here.

[0093] Step 507: If the second probability is greater than or equal to the preset probability threshold, determine that the file to be processed is a malicious inducement file; otherwise, determine that the file to be processed is not a malicious inducement file.

[0094] In this step, the preset probability threshold may be 0.7, that is, only when the second probability is greater than or equal to 0.7, the file to be processed is considered to be a malicious inducement file, otherwise, it is determined that the file to be processed is not a malicious inducement file. That is, in this embodiment, the classifier can ultimately output the classification result of whether the file to be processed is a malicious inducement file or a non-malicious inducement file.

[0095] The above-mentioned malicious decoy file detection method simultaneously performs image feature classification and text content recognition on the homepage preview image of the file containing macro code, and combines the two types of results to determine whether the file contains macro decoy behavior. Among them, the above-mentioned image feature classification and text content recognition can be implemented through the corresponding model, so that based on the recognition and classification results of the model, it can be accurately determined whether the file is a macro decoy file, thereby effectively improving the recognition accuracy. The file types to be processed referred to in this method include but are not limited to Office files. In fact, all files that meet the characteristics of malicious decoy files can be applied to this method. Its advantages can be found in the above Figure 4A The description of the corresponding embodiment will not be repeated here!

[0096] Please see Figure 6 , which is a malicious inducement file detection device 600 of an embodiment of the present application, and the device can be applied to Figure 1 The electronic device 1 shown can be applied to Figure 2A In the malicious file detection scenario shown in the figure, the detection accuracy of malicious files and the versatility of the detection method are improved. The device includes: a first acquisition module 601, an identification module 602 and a determination module 603, and the principle relationship of each module is as follows:

[0097] The first acquisition module 601 is used to acquire image information of the file to be processed.

[0098] The recognition module 602 is used to perform image processing on the image information to obtain the image processing result of the image information, and perform text recognition on the image information to obtain the text features in the image information.

[0099] The determination module 603 is used to determine whether the file to be processed is a malicious inducement file according to the image processing result and the text feature.

[0100] In one embodiment, the recognition module 602 is used to: input the image information into a preset recognition model, and output an image processing result of the image information, wherein the preset recognition model is at least used to extract image features from the image information.

[0101] In one embodiment, it further includes: a second acquisition module 604, which is used to acquire a sample file data set, select a training set and a test set from the sample file data set, and each sample file in the training set and the test set is labeled with a label of whether it is a malicious inducement file. A building module 605 is used to train a neural network model using the training set, and test the trained model using the test set to obtain a preset recognition model.

[0102] In one embodiment, the establishment module 605 is used to: train the neural network model with the training set to obtain a primary classification model. Use the test set to test the primary classification model, collect the wrong sample set of the primary classification model on the test set, and the recognition results of the samples in the wrong sample set are different from the labels of the corresponding samples in the test set. Select similar samples whose similarity with the wrong sample reaches a first threshold from the remaining sample file data set, and the remaining sample file data set is the data set after the training set is removed from the sample file data set. Add similar samples to the training set, and use the updated training set to train the neural network model, iteratively update the training set, until a preset recognition model is established whose test results reach a preset accuracy rate.

[0103] In one embodiment, the recognition module 602 is further used to: perform text recognition on the image information to obtain the text content in the image information as the text feature in the image information. And / or, perform text recognition on the image information to obtain the text content in the image information; extract the word vector of the text content, and based on the word vector, extract the semantic feature of the text content to obtain the text feature in the image information.

[0104] In one embodiment, the image processing result includes a first probability that the file to be processed is a malicious inducement file. The text feature includes the text content in the image information. The determination module 603 is used to: determine whether there are identification words for malicious inducement files in the text content based on the text features. If there are identification words in the text content, determine that the file to be processed is a malicious inducement file. If there are no identification words in the text content, determine whether the first probability is greater than or equal to a preset probability threshold. If the first probability is greater than or equal to the preset probability threshold, determine that the file to be processed is a malicious inducement file, otherwise, determine that the file to be processed is not a malicious inducement file.

[0105] In one embodiment, the image processing result includes the image features of the file to be processed; the determination module 603 is used to: fuse the image processing result and the text features to generate the fused features of the file to be processed. According to the fused features, a second probability that the file to be processed is a malicious inducement file is determined. If the second probability is greater than or equal to a preset probability threshold, the file to be processed is determined to be a malicious inducement file; otherwise, the file to be processed is determined not to be a malicious inducement file.

[0106] For a detailed description of the malicious file detection device 600, please refer to the description of the relevant method steps in the above embodiment.

[0107] The embodiment of the present invention also provides a non-transitory electronic device readable storage medium, including: a program, when it is run on an electronic device, the electronic device can execute all or part of the process of the method in the above embodiment. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (Hard Disk Drive, abbreviated: HDD) or a solid-state drive (SSD), etc. The storage medium can also include a combination of the above types of memory.

[0108] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for detecting malicious files. It is characterized in that include: Get the image information of the file to be processed; Performing image processing on the image information to obtain an image processing result of the image information, and performing text recognition on the image information to obtain text features in the image information; Determining whether the file to be processed is a malicious inducement file according to the image processing result and the text feature; The image processing result is obtained by processing with a preset recognition model, and the preset recognition model is obtained in the following manner: obtaining a sample file data set, selecting a training set and a test set from the sample file data set, wherein each sample file in the training set and the test set is labeled with a label of whether it is a malicious inducement file; using the training set to train a neural network model to obtain a primary classification model; using the test set to test the primary classification model, collecting a set of false example samples of the primary classification model on the test set, wherein the recognition results of the samples in the false example sample set are different from the labels of the corresponding samples in the test set; Select similar samples whose similarity with the wrong example samples reaches a first threshold from the remaining sample file data set, and the remaining sample file data set is the data set after the sample file data set is removed from the training set; add the similar samples to the training set, and use the updated training set to train the neural network model, and iteratively update the training set until the preset recognition model with a test result reaching a preset accuracy rate is established.

2. The method according to claim 1, It is characterized in that The performing image processing on the image information to obtain an image processing result of the image information includes: The image information is input into a preset recognition model, and the image processing result of the image information is output, wherein the preset recognition model is at least used to extract image features from the image information.

3. The method according to claim 1, It is characterized in that The performing text recognition on the image information to obtain text features in the image information includes: Performing text recognition on the image information to obtain text content in the image information as the text feature in the image information; and / or, performing text recognition on the image information to obtain text content in the image information; The word vector of the text content is extracted, and based on the word vector, the semantic features of the text content are extracted to obtain the text features in the image information.

4. The method according to claim 1, It is characterized in that The image processing result includes a first probability that the file to be processed is a malicious inducement file; the text feature includes the text content in the image information; The determining, based on the image processing result and the text feature, whether the file to be processed is a malicious inducement file includes: Based on the text features, determining whether there are any identification words for malicious files in the text content; If the identification word exists in the text content, determining that the file to be processed is a malicious inducement file; If the identification word does not exist in the text content, determining whether the first probability is greater than or equal to a preset probability threshold; If the first probability is greater than or equal to the preset probability threshold, it is determined that the file to be processed is a malicious inducement file; otherwise, it is determined that the file to be processed is not a malicious inducement file.

5. The method according to claim 1, It is characterized in that The image processing result includes the image features of the file to be processed; and determining whether the file to be processed is a malicious inducement file according to the image processing result and the text features includes: Fusing the image features and the text features to generate fused features of the file to be processed; Determining, based on the fusion feature, a second probability that the file to be processed is a malicious inducement file; If the second probability is greater than or equal to a preset probability threshold, it is determined that the file to be processed is a malicious inducement file; otherwise, it is determined that the file to be processed is not a malicious inducement file.

6. A malicious inducement file detection device, It is characterized in that include: A first acquisition module, used to acquire image information of a file to be processed; A recognition module, used to perform image processing on the image information to obtain an image processing result of the image information, and perform text recognition on the image information to obtain text features in the image information; A determination module, used to determine whether the file to be processed is a malicious inducement file according to the image processing result and the text feature; The image processing result is obtained by processing with a preset recognition model, and the preset recognition model is obtained in the following manner: obtaining a sample file data set, selecting a training set and a test set from the sample file data set, wherein each sample file in the training set and the test set is labeled with a label of whether it is a malicious inducement file; using the training set to train a neural network model to obtain a primary classification model; using the test set to test the primary classification model, collecting a set of false example samples of the primary classification model on the test set, wherein the recognition results of the samples in the false example sample set are different from the labels of the corresponding samples in the test set; Select similar samples whose similarity with the wrong example samples reaches a first threshold from the remaining sample file data set, and the remaining sample file data set is the data set after the sample file data set is removed from the training set; add the similar samples to the training set, and use the updated training set to train the neural network model, and iteratively update the training set until the preset recognition model with a test result reaching a preset accuracy rate is established.

7. An electronic device, It is characterized in that include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 5.

8. A non-transitory electronic device readable storage medium, It is characterized in that The invention comprises: a program, which, when executed by an electronic device, causes the electronic device to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Model optimization method and device, and electronic equipment

    CN111027707A

  • Method and system for identifying rogue behavior of Android application

    CN111832021A