A Malicious Application Detection Method and Device Based on Randomness
Through the randomness-based malicious application detection method, key features of the application are extracted and analyzed, word randomness is calculated and inputted into AI model for detection, the problem of slow update of malicious feature library in the existing technology is solved, and higher detection accuracy and timeliness are achieved.
Patent Information
- Application Number
- CN202010721432.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-07-24
AI Technical Summary
When detecting malicious applications, existing mobile security software relies on the update of the malicious feature library and has weak ability to learn new malicious detection, resulting in the inability to effectively protect new malicious applications.
A malicious application detection method based on randomness is adopted. By analyzing the target files in the application installation package, key features are extracted, and the word randomness of the key features is calculated, the randomness data is spliced into a digital feature vector, and the trained AI model is input for malicious detection.
It improves the accuracy and timeliness of malicious program detection, and solves the problems of difficulty in extracting rules, low coverage, poor scalability, and easy to be bypassed in traditional methods.
Smart Images

Figure CN113971281B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of mobile network security, and in particular, to a method and device for detecting malicious application programs based on randomness. Background Art
[0002] Currently, many security vendors have also entered the mobile security field. The basic principle of these software for antivirus is to confirm intrusion behavior by matching known malicious trojan features and perform active defense in ways such as firewalls and dynamic monitoring. However, the disadvantage is that it relies on the update of the malicious feature library and has a weak ability to learn new malicious detections.
[0003] However, new malicious application programs emerge in an endless stream, and their maliciousness varies. Relying on the malicious feature library for malicious detection, it is difficult to achieve an ideal security protection effect due to the untimely update of the malicious feature library. Summary of the Invention
[0004] In view of the problems existing in the prior art, embodiments of the present invention provide a method and device for detecting malicious application programs based on randomness. On the basis of analyzing a large number of mainstream malicious software on the Android platform, the attack intentions and means of malicious software on the Android platform are summarized, and the AI detection of the maliciousness of application programs is realized through a deep learning algorithm, providing a new direction for the malicious detection of application programs on the Android platform.
[0005] In a first aspect, embodiments of the present invention provide a method for detecting malicious application programs based on randomness, including:
[0006] Analyze the target files in the installation package of the application program to be detected, and extract the key features in the target files. The form of the key features is a string, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension;
[0007] Obtain the word randomness of the key features and splice the word randomness into a digital feature vector;
[0008] Input the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.
[0009] Further, the obtaining the word randomness of the key features and splicing the word randomness into a digital feature vector specifically includes:
[0010] Obtain the adjacent letter frequencies of any two adjacent letters of the key feature string in alphabetical order of the key feature string in sequence;
[0011] Obtain the word randomness of the key feature based on the adjacent letter frequencies of any two adjacent letters;
[0012] Concatenate the word randomness into a digital feature vector;
[0013] Wherein, the adjacent letter frequency is to regard the string as a word of a preset language text, and obtain the adjacent frequency between two letters through the word randomness calculation rule of the preset language text; the word randomness is obtained through the adjacent letter frequency.
[0014] Further, the concatenating the word randomness into a digital feature vector specifically includes:
[0015] Based on each randomness data in the English word randomness and / or the Chinese pinyin randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, convert it into an encoded number composed of 0 and 1;
[0016] Concatenate all the encoded numbers in a preset order into a digital feature vector composed of 0 and 1.
[0017] Further, the concatenating the word randomness into a digital feature vector specifically includes:
[0018] Based on each randomness data in the word randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, convert it into an encoded number composed of 0 and 1; concatenate all the encoded numbers in a preset order into a digital feature vector composed of 0 and 1.
[0019] Further, the specifically including that after performing feature transformation on each randomness data respectively, through one-hot encoding, convert it into an encoded number composed of 0 and 1 specifically includes:
[0020] Based on all the randomness data, obtain N numerical segments according to the first preset rule, sort the N numerical segments in ascending order of numerical value, and obtain the sorting position of each numerical segment, where N is an integer greater than 0;
[0021] Match each randomness data with the N numerical segments according to the second preset rule, so that each randomness data matches a numerical segment, and use the sorting position of the matched numerical segment as the feature transformation number of each randomness data;
[0022] Perform one-hot encoding on the feature transformation numbers of each randomness data, so as to obtain an encoded number composed of 0 and 1.
[0023] Further, the network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method:
[0024] Obtain samples of application installation packages and the malicious labels of the application installation package samples;
[0025] Parse the second target file in the application installation package sample, extract the second key features in the second target file, the form of the key features is a string, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension;
[0026] Obtain the word randomness of the second key feature, and splice the word randomness into a second digital feature vector;
[0027] Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.
[0028] Further, the inputting the second digital feature vector and the malicious label of the application installation package sample into the built AI model and training the AI model specifically includes:
[0029] Convert the malicious label of the application installation package sample into the malicious label of the second digital feature vector;
[0030] Input the second digital feature vector and the malicious label of the second digital feature vector into the built AI model, and train the AI model.
[0031] In a second aspect, an embodiment of the present invention provides an electronic device, including:
[0032] At least one processor; and
[0033] At least one memory communicatively connected to the processor, wherein:
[0034] The memory stores program instructions executable by the processor, and the processor can execute the method for detecting malicious application programs based on randomness and any optional embodiment thereof in the first aspect of the embodiments of the present invention by calling the program instructions.
[0035] In a third aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions execute the method for detecting malicious application programs based on randomness and any optional embodiment thereof in the first aspect of the embodiments of the present invention.
[0036] The malicious application detection method based on randomness provided by the embodiments of the present invention extracts the key feature strings of the target files in the installation package of the application to be detected, obtains the word randomness corresponding to the key feature strings, and splices all the obtained randomness data into a digital feature vector; based on the trained AI model, the AI model performs operations on the digital feature vector, and the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application can be obtained. The embodiments of the present invention solve the problems existing in the traditional malicious application detection based on manual extraction rules, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed, and have higher accuracy and timeliness for malicious program detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic flowchart of the malicious application detection method based on randomness according to the embodiments of the present invention;
[0039] Figure 2 It is a schematic diagram of numerical segmentation according to the embodiments of the present invention;
[0040] Figure 3 It is a schematic diagram of the network structure of the AI model according to the embodiments of the present invention;
[0041] Figure 4 It is a schematic flowchart of the AI model training according to the embodiments of the present invention;
[0042] Figure 5 It is a malicious application detection device based on randomness according to the embodiments of the present invention;
[0043] Figure 6 It is a schematic framework diagram of an electronic device according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0045] In view of the problems of the prior art, in an embodiment of the present invention, an algorithm is redesigned based on the DeepFM algorithm in the deep learning method to obtain an Artificial Intelligence (AI) model; based on the extraction of key features from a large number of Android platform application programs, the AI model is trained with the extracted key features, and finally an AI model for detecting malicious application programs on the Android platform is obtained, where the training process includes: extracting key features from sample data (application program installation package samples); performing feature transformation on the extracted key features to obtain digital feature vectors corresponding to the sample data; using the digital feature vectors corresponding to the sample data and the malicious labels corresponding to the sample data as input data of the AI model to train the AI model, and evaluating the effect through a model evaluation strategy to obtain an AI model that meets the expected requirements.
[0046] When detecting malicious applications, key features are extracted from the target files in the application installation package to be detected, the word randomness of the key features is calculated (further feature transformation is performed), and they are concatenated into a digital feature vector corresponding to the application installation package to be detected; the digital feature vector corresponding to the application installation package to be detected is used as the input data of the AI model, and after the AI model performs operations, it outputs the probability that the application installation package to be detected is a malicious application and / or the probability that it is a non-malicious application.
[0047] In an embodiment of the present invention, both the AI model training stage and the malicious application program detection stage include key feature extraction and calculation of the word randomness of key features (further feature transformation), and their processing methods are exactly the same; in all alternative solutions of the embodiment of the present invention, the solution adopted in the malicious application detection stage of the application program is consistent with the solution adopted in the AI model training stage to obtain the expected detection effect.
[0048] Next, from the perspective of the malicious application program detection stage, the randomness-based malicious application program detection method described in the embodiment of the present invention will be described in detail.
[0049] Figure 1 It is a schematic flowchart of the randomness-based malicious application program detection method described in an embodiment of the present invention. As Figure 1 shown, the randomness-based malicious application program detection method includes:
[0050] 101. Analyze the target files in the application installation package to be detected, extract the key features in the target files, the form of the key features is a string, and the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;
[0051] 102. Obtain the word randomness of the key features and splice the word randomness into a digital feature vector;
[0052] 103. Input the digital feature vector into the trained AI model to obtain the malicious detection result of the application to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.
[0053] When the embodiment of the present invention performs malicious detection on an application, first, an application installation package to be detected is obtained, the target file in the application installation package to be detected is parsed, and the key features in the target file are extracted. The key features include at least one dimension information among the behavior dimension, the permission dimension, and the content dimension. Each dimension information may include one or more different pieces of information. For example, the behavior dimension may include Behavior 1, Behavior 2, Behavior 3, etc., the permission dimension may include Permission 1, Permission 2, Permission 3, Permission 4, etc., and the content dimension may include Content 1, Content 2, Content 3, etc. The forms of all the obtained key features are strings.
[0054] Before performing step 101, the embodiment of the present invention further includes basic identification and verification of the application installation package, and the main purpose is to identify damaged and abnormally formatted installation packages. The magic header is the magic number at the beginning of a file used to represent different file types. The embodiment of the present invention performs file identification and verification through the magic header, including:
[0055] File validity verification: Basic parsing and verification of the tail of the Zip file, which is to check whether the magic header is 0x06054b50; verification of each sub-file entry, checking whether the magic header of each entry is 0x02014b50;
[0056] Integrity verification of key sub-files: Format verification of certificates, and the verification method is the verification of the standard X509 certificate chain; verification of the classes.dex file, checking whether the file header is "dex\n035", "dex\n036", "dex\n037", "dex\n038", etc., and the above strings also correspond to the specific versions of the classes.dex file; verification of the AndroidManifest.xml file, mainly verifying whether the main magic header is 0x00080003, the magic header of the string area is 0x001c0001, and the magic header of the resource area is 0x00080180; verification of the resources.arsc file, mainly verifying whether its main magic header is 0x0002.
[0057] The verification of the above key sub-files may further include: whether the file size declared in the file header field is consistent with its size in the zip package.
[0058] In step 102, the word randomness of the extracted key features is calculated. The implementation of the word randomness in the embodiments of the present invention may include the word randomness of one or more languages, such as English word randomness, Japanese word randomness, French word randomness, Chinese pinyin randomness, and so on. In application, one type of word randomness can be used according to the application environment, or multiple types of word randomness can be used simultaneously.
[0059] For example, for one or more extracted key features, the English word randomness of each key feature can be calculated, and all the English word randomness values can be concatenated into a digital feature vector; or the Chinese pinyin randomness of each key feature can be calculated, and all the Chinese pinyin randomness values can be concatenated into a digital feature vector; or the English word randomness and Chinese pinyin randomness of each key feature can be calculated simultaneously, and all the English word randomness and Chinese pinyin randomness values can be concatenated into a digital feature vector. The embodiments of the present invention do not limit the specific concatenation order, as long as the concatenation order in the malicious application detection stage is consistent with that in the AI model training stage.
[0060] In the analysis of a large number of application program samples, it is found that many malicious samples containing viruses are automatically generated based on special tools. Since such samples are automatically generated, the keywords in string form such as package names, program names, and developer information often have strong randomness. Therefore, analyzing the randomness of application program keywords is a very effective virus detection method. Randomness is a determination of whether a word is a reasonable word or a group of randomly combined letters. The closer a word is to a common word, the lower its randomness; conversely, the higher its randomness, and the higher the randomness, the greater the possibility of a malicious application.
[0061] The embodiments of the present invention directly form the randomness of the key features into a digital feature vector as the input data of the AI model, and judge malicious applications through the randomness of the application program key features.
[0062] In step 103, the digital feature vector is input into a pre-trained AI model, and the maliciousness detection result of the application program to be detected can be obtained through calculation; the maliciousness detection result in the embodiments of the present invention refers to the probability that the application program to be detected is a malicious application and / or the probability of a non-malicious application. The AI model in the embodiments of the present invention can output the probability that the application program is a malicious application, or output the probability that the application program is a non-malicious application, or output the probability that the application program is a malicious application and the probability of a non-malicious application simultaneously.
[0063] Suppose the AI model outputs the probabilities of malicious applications and non-malicious applications simultaneously. If the former is greater than the latter, the application is non-malicious; otherwise, the application is malicious. For example, if it outputs "[0.98654, 0.01346]", where 0.98654 is the probability of a non-malicious application and 0.01346 is the probability of a malicious application, it means this program is non-malicious, thus completing the determination of the maliciousness of the application.
[0064] In the malicious application detection method based on randomness according to the embodiments of the present invention, key features of the target file in the installation package of the application to be detected are extracted, the word randomness corresponding to the key feature string is obtained, and all the obtained randomness data is spliced into a digital feature vector; based on the already trained AI model, by operating on the digital feature vector through the AI model, the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application can be obtained. The embodiments of the present invention solve the problems existing in the traditional malicious application detection based on manually extracting rules, such as difficult rule extraction, low coverage, poor scalability, and being easily bypassed, and have higher accuracy and timeliness for malicious program detection.
[0065] Based on the above embodiments, step 101 of parsing the target file in the installation package of the application to be detected and extracting the key features in the target file specifically includes:
[0066] 101.1, parse the information of each sub-file in the installation package of the application to be detected to obtain the target file;
[0067] 101.2, in the target file, extract the target string that matches a specific keyword, which is the key feature; the target string includes at least one dimension information of the behavior dimension, permission dimension, and content dimension;
[0068] In step 101.1 of the embodiments of the present invention, the target file is obtained from the sub-files of the installation package of the application to be detected, and the target file can be files such as AndroidManifest.xml and classes.dex in the sub-files.
[0069] In the embodiments of the present invention, the key features in the target file are extracted by means of keyword matching. The specific keyword in step 101.2 is a keyword related to behavior, permission, and content, and the target string matched by the specific keyword is the string information related to behavior, permission, and content. By matching all the specific keywords set in the embodiments of the present invention, all the information of the behavior dimension, permission dimension, and content dimension may be matched, or only part of the information may be matched.
[0070] Specifically, the behavior dimension includes the behavior information when the application is running;
[0071] The permission dimension includes permission information required for an application to perform specific actions;
[0072] The content dimension includes at least one of the following information: application package name, program name, developer information, and language information.
[0073] The key features described in the embodiments of the present invention include at least one dimension information among the behavior dimension, permission dimension, and content dimension; specifically, the extraction methods for each dimension information include but are not limited to the following descriptions, and each dimension information includes at least one of the descriptions:
[0074] In the behavior dimension, the behavior of an application refers to actions such as accessing local files, sending text messages, connecting to the network, etc. performed during program operation, and specifically manifests as the existence of such behavior information in the classes.dex file in the form of specific class names, method names, and variable names. The behavior of an application is an important basis for virus detection, and many virus-infected applications will have several high-risk behaviors. Analyzing the list of high-risk behaviors of an application is an important basis for determining whether the application carries a virus. Based on the analysis of a large number of existing applications, the relationship between the above key symbols and the application behavior can be found, and by corresponding the specific combination of key symbols with the application behavior, a behavior feature library can be formed.
[0075] For example:
[0076] (1) The behavior "SMS_Send", indicating whether there is a behavior of sending text messages, and the corresponding behavior rules are one of the following "class -> method":
[0077] android / telephony / SmsManager->sendMultipartTextMessage;
[0078] android / telephony / SmsManager->sendTextMessage;
[0079] android / telephony / gsm / SmsManager->sendTextMessage;
[0080] android / telephony / gsm / SmsManager->sendMultipartTextMessage;
[0081] android / telephony / SmsManager->sendDataMessage;
[0082] android / telephony / gsm / SmsManager->sendDataMessage;
[0083] (2) The behavior "SMS_Scan" represents the behavior of SMS search, and the corresponding behavior rule is to include the "class -> method" combination:
[0084] android / database / Cursor->moveToFirst, and the static string contains the url: content: / / sms;
[0085] Or include the "class -> method" combination:
[0086] android / database / Cursor->moveToPosition, and the static string contains the url: content: / / sms;
[0087] In the dimension of permissions, when an application performs a file reading behavior, it needs to declare the permission to read the file first. Some of these permissions are sensitive permissions. Therefore, whether certain permissions are declared or not has certain reference value for judging the maliciousness of the application. Preferably, the embodiments of the present invention obtain information about permission declarations from the "<uses-permission name = "xxx">" field of the AndroidManifest.xml file.
[0088] In the dimension of content, the embodiments of the present invention parse the AndroidManifest.xml file. After decryption, this file is a standard xml file. Query its "<manifest package = "xxx">" field, where xxx is the package name of the application; query the "<application lable = "xxx">" field, where xxx represents the index of the application name in the resources.arsc file. Parsing the resources.arsc can obtain the application name through this index; the certificate file is located in the META-INF directory and is in the standard X509 format. Calling the relevant openssl interfaces can parse the certificate to obtain the developer information of the application. openssl is a large encryption and decryption library for common standard formats. Since the certificate belongs to the X509 standard format, directly calling the relevant interfaces for parsing can obtain the developer information.
[0089] Based on the analysis of a large number of existing applications, the embodiments of the present invention extract information in dimensions such as specific behaviors, permissions, sensitive strings, and certificates included in the application installation package. Malicious applications are highly distinguishable from normal applications in these dimensions, which is an important basis for judging the maliciousness of applications.
[0090] Based on any of the above optional embodiments, in step 102, obtaining the word randomness of the key feature and concatenating the word randomness into a digital feature vector specifically includes:
[0091] 102.1, according to the alphabetical order of the key feature string, sequentially obtain the letter adjacent frequency of any two adjacent letters of the key feature string;
[0092] 102.2, based on the letter adjacent frequency of any two adjacent letters, obtain the word randomness of the key feature;
[0093] 102.3, concatenate the word randomness into a digital feature vector;
[0094] Wherein, the letter adjacent frequency is to regard the string as a word of a preset language text, and obtain the adjacent frequency between two letters through the word randomness calculation rule of the preset language text; the word randomness is obtained through the letter adjacent frequency.
[0095] For example, if the preset language is English, the letter adjacent frequency is to regard the string as English and obtain the adjacent frequency between two letters through the English word randomness calculation rule;
[0096] If the preset language is Chinese, the letter adjacent frequency is to regard the string as Chinese pinyin and obtain the adjacent frequency between two letters through the Chinese pinyin randomness calculation rule;
[0097] Correspondingly, the English word randomness is obtained through the letter adjacent frequency of the English word; the Chinese pinyin randomness is obtained through the letter adjacent frequency of the Chinese pinyin.
[0098] In the embodiments of the present invention, the word randomness calculation rule is obtained by analyzing the alphabetical order of a large number of natural languages. The English word randomness calculation rule is obtained by analyzing the alphabetical order of English words in a large number of natural languages. The Chinese pinyin randomness calculation rule is obtained by analyzing the alphabetical order of Chinese pinyin in a large number of natural languages.
[0099] For an application program, the string randomness calculation can use any set of strings parsed from the installation package. By calculating the randomness of the string, it can be determined whether a string is a reasonable word or a group of random letter combinations, so as to determine whether the corresponding application program is malicious or non-malicious.
[0100] In practical applications, there are many rules for calculating randomness. The string rules in different application fields also have specific meanings. Therefore, the randomness calculation rules in different application fields are also different. The embodiments of the present invention do not limit this. The key point applied in the embodiments of the present invention is that regardless of the rule used to calculate the randomness of the strings in this field, the lower the randomness, the greater the possibility that the string is a reasonable word; the higher the randomness, the greater the possibility that the string is a random combination of letters, and this is used as the basis for judging whether the application program is malicious or non-malicious.
[0101] By analyzing the alphabetical order of the strings in different languages and scripts, the adjacent letter frequencies of any two letters can be obtained, and these adjacent letter frequencies can be pre-composed into a frequency library file. According to the alphabetical order of the key feature string to be calculated in the embodiments of the present invention, the adjacent letter frequencies of the adjacent letters are queried in turn, and then the occurrence frequencies of all adjacent letters are combined to obtain a value of the overall adjacent frequency of a string, named "entropy". The lower the entropy of the string, the higher the randomness of the string, indicating that the string is less like a regular word.
[0102] The calculation method of the entropy of the string in the embodiments of the present invention is as follows:
[0103] (1) Let entropy represent entropy, and set the initial entropy to 0;
[0104] (2) Traverse the string. For every two adjacent letters of the string, query the adjacent letter frequency of this two-letter combination in the frequency library file, and record it as prob;
[0105] (3) Update entropy according to the following formula:
[0106] entropy = entropy - prob * log 2 prob;
[0107] (4) Repeat steps (2) and (3) until the string traversal ends.
[0108] The following uses English words and Chinese pinyin to illustrate randomness.
[0109] Suppose the package name of application program 1 is: com.iizirgui.vcpnocuc;
[0110] The developer information is:
[0111] CN = hmqmawavzg.oaa, OU = tdbkusvGGE, O = ppCIFivngI,
[0112] L = fRitORayrU, ST = DtTJWEpbqs, C = QL,
[0113] Taking the package name "com.iizirgui.vcpnocuc" as English, the adjacent frequency between every two letters is obtained according to the English word randomness calculation rule, and the English entropy of its package name is 0.160173 according to the calculation method of the entropy of the string; taking the package name "com.iizirgui.vcpnocuc" as Chinese pinyin, the adjacent frequency between every two letters is obtained according to the Chinese pinyin randomness calculation rule, and the Chinese pinyin entropy is 0.118652 according to the calculation method of the entropy of the string;
[0114] Similarly, the English entropy of the CN segment of the developer information is 0.117621, and the Chinese pinyin entropy is 0.093654; the English entropy of the OU segment of the developer information is 0.130641, and the Chinese pinyin entropy is 0.113263; the English entropy of the O segment of the developer information is 0.222635, and the Chinese pinyin entropy is 0.099876.
[0115] The above examples obtain the English word randomness and the Chinese pinyin randomness at the same time, and splice the English word randomness and the Chinese pinyin randomness into a digital feature vector according to a preset order; as mentioned above, the embodiment of the present invention does not limit the splicing order, and one of the splicings is as follows:
[0116] [0.160173, 0.118652, 0.117621, 0.093654, 0.130641, 0.113263, 0.222635, 0.099876].
[0117] In the above embodiment, for the extracted key features (at least one dimension information among the behavior dimension, the permission dimension, and the content dimension), after calculating the randomness, the obtained randomness data is directly spliced into a digital feature vector and input into the trained AI model for operation, and the maliciousness detection result of the corresponding application to be detected can be obtained.
[0118] Based on any of the above optional embodiments, in step 102.3, splicing the word randomness into a digital feature vector specifically includes:
[0119] Based on each randomness data in the word randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, it is converted into an encoded number composed of 0 and 1; all the encoded numbers are spliced into a digital feature vector composed of 0 and 1 according to a preset order.
[0120] In the embodiments of the present invention, the randomness data is subjected to feature transformation and one-hot encoding. If multiple types of word randomness are used, the randomness data of each type of word can be processed separately, that is, feature transformation and one-hot encoding are respectively performed based on the randomness data of each type of word, and finally all the encoded numbers are concatenated in a preset order into a digital feature vector composed of 0s and 1s.
[0121] Preferably, in the embodiments of the present invention, for each randomness data in the English word randomness and / or Chinese Pinyin randomness, after each randomness data is subjected to feature transformation respectively, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s; all the encoded numbers are concatenated in a preset order into a digital feature vector composed of 0s and 1s.
[0122] This embodiment is divided into three types according to the calculation of randomness: calculation according to the English word randomness, calculation according to the Chinese Pinyin randomness, and calculation according to both the English word randomness and the Chinese Pinyin randomness simultaneously.
[0123] In the embodiments of the present invention, the calculated randomness data can also be subjected to feature transformation and then concatenated into a digital feature vector; according to the three ways of calculating randomness, the feature transformation of the randomness data also includes three types:
[0124] (1) For each randomness data in the English word randomness, after each randomness data is subjected to feature transformation respectively, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s; all the encoded numbers are concatenated in a preset order into a digital feature vector composed of 0s and 1s.
[0125] (2) For each randomness data in the Chinese Pinyin randomness, after each randomness data is subjected to feature transformation respectively, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s; all the encoded numbers are concatenated in a preset order into a digital feature vector composed of 0s and 1s.
[0126] (3) For each randomness data in the English word randomness, after each randomness data is subjected to feature transformation respectively, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s; for each randomness data in the Chinese Pinyin randomness, after each randomness data is subjected to feature transformation respectively, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s; all the encoded numbers after the conversion of the English word randomness and the Chinese Pinyin randomness are concatenated in a preset order into a digital feature vector composed of 0s and 1s.
[0127] (1) and (2) perform feature transformation on the randomness of English words and the randomness of Chinese pinyin respectively, (3) performs feature transformation on the randomness of English words and the randomness of Chinese pinyin simultaneously. Each randomness data will obtain a feature transformation number, and then the feature transformation numbers are one-hot encoded, and finally concatenated into a digital feature vector composed of 0s and 1s. Through feature transformation, while retaining more feature information, the generalization ability of the AI model can be improved.
[0128] Similarly, the embodiments of the present invention do not limit the concatenation order of the digital feature vectors after feature transformation and one-hot encoding, as long as the concatenation order in the malicious application detection stage is consistent with that in the AI model training stage.
[0129] One-hot encoding is also called one-hot encoding. The method is to use an N-bit status register to encode N states. Each state has its own independent register bit, and at any time, only one of them is valid.
[0130] For example, encoding six states:
[0131] The natural sequence code is: 000, 001, 010, 011, 100, 101;
[0132] The one-hot encoding is: 000001, 000010, 000100, 001000, 010000, 100000.
[0133] The embodiments of the present invention have flexible and diverse ways of performing feature transformation on randomness data. Based on the specific data of the extracted key features and the operation requirements of different scenarios, different feature transformation methods can be adopted.
[0134] Based on any of the above optional embodiments, after respectively performing feature transformation on each randomness data, through one-hot encoding, it is converted into an encoded number composed of 0s and 1s, specifically including:
[0135] Based on all randomness data, obtain N numerical segments according to the first preset rule, and sort the N numerical segments according to the numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0;
[0136] Match each randomness data with the N numerical segments according to the second preset rule, so that each randomness data matches a numerical segment, and use the sorting position of the matched numerical segment as the feature transformation number of each randomness data;
[0137] Perform one-hot encoding on the feature transformation number of each randomness data, so as to obtain an encoded number composed of 0s and 1s.
[0138] Figure 2This is a schematic diagram of numerical segmentation for an embodiment of the present invention. Assume that there are a total of 1150 randomness data obtained after calculating the key features extracted in the embodiment of the present invention, that is, 1150 randomness data are obtained. Then, segmentation is performed according to the total of 1150 randomness data, and the segmentation method is the first preset rule. The first preset rule can be to segment according to the numerical magnitudes of the 1150 randomness data, or to segment according to the number of the 1150 randomness data, or to segment according to a mixture of numerical magnitudes and the number of data. The specific segmentation method can be determined according to requirements, and the embodiment of the present invention does not limit this.
[0139] As Figure 2 shown, assume that in the embodiment of the present invention, through the first preset rule, after segmenting the 1150 randomness data in the example, N numerical segments are obtained. After sorting these N numerical segments by size, the sorting positions are 1, 2,..., N respectively.
[0140] After sorting the numerical segments, the embodiment of the present invention "assigns seats" to all the randomness data according to the second preset rule, where the second preset rule is a matching method corresponding to the first preset rule, that is: if the first preset rule is to segment according to numerical magnitudes, then the second preset rule matches according to numerical magnitudes; if the first preset rule is to segment according to the number of data, then the second preset rule matches according to the number ranking of the data; if the first preset rule is to segment according to a mixture of numerical magnitudes and the number of data, then the second preset rule matches according to a mixture of numerical magnitudes and the number of data.
[0141] Please refer to Figure 2 , taking the first preset rule of segmenting according to numerical magnitudes as an example, for the 1150 randomness data in the above example, the 1150 randomness data are respectively matched with N numerical segments. If the numerical values of 50 of the data are within the numerical range of segment 1, then the sorting position 1 is used as the feature transformation number for these 50 data, that is, each of these 50 data obtains the feature transformation number 1; if the numerical values of 100 of the data are within the numerical range of segment 2, then the sorting position 2 is used as the feature transformation number for these 100 data, that is, each of these 100 data obtains the feature transformation number 2; and so on. Each randomness data will obtain a feature transformation number, which will not be elaborated further.
[0142] After segmenting and matching each randomness data, each randomness data obtains a feature transformation number. The embodiment of the present invention performs one-hot encoding on each feature transformation number to obtain an encoded number composed of 0 and 1. The one-hot encoding process in the following embodiments is the same and will not be elaborated further hereafter.
[0143] It should be noted that since the word randomness in the embodiments of the present invention can include the word randomness of multiple languages, such as English word randomness, Chinese pinyin randomness, etc., the feature transformation and one-hot encoding can be performed on the English word randomness, Chinese pinyin randomness, and the first numerical data respectively. The English word randomness and Chinese pinyin randomness consist of different key feature strings of English word randomness and Chinese pinyin randomness, and the first numerical data also includes data with different key features. Therefore, in actual implementation, it can be processed according to different key features respectively, that is, the data of one feature transformation and one-hot encoding belongs to one key feature. For example, all of the above 1150 data are the quantity of metadata, or all of the 1150 data are the English word randomness calculated from the application package name, or all of the 1150 data are the Chinese pinyin randomness calculated from the application package name, and so on. Performing feature transformation and one-hot encoding according to the same key feature can generalize the data of this key feature, thereby improving the generalization ability of the AI model. The following feature transformations can all apply this strategy and will not be elaborated hereafter.
[0144] In the embodiments of the present invention, feature transformation is performed on the randomness data, which can classify and discretize the randomness data with similar characteristics (such as numerical values that are different but very close), and can improve the generalization ability of the AI model during the training of the AI model. Preferably, the embodiments of the present invention provide three feature transformation methods, which are the feature transformation methods according to numerical values, according to the number of data, and according to the mixture of numerical values and the number of data, as follows:
[0145] The first feature transformation method: The first preset rule is to segment according to the numerical size, and the second preset rule is to match according to the numerical size;
[0146] The second feature transformation method: The first preset rule is to segment according to the number of data, and the second preset rule is to match according to the serial number of the number of data;
[0147] The third feature transformation method: The first preset rule is to segment according to the mixture of numerical size and the number of data, and the second preset rule is to match according to the mixture of numerical size and the number of data.
[0148] The following will describe each feature transformation method in detail.
[0149] Based on any of the above optional embodiments, after each randomness data is respectively subjected to feature transformation, through one-hot encoding, it is converted into encoded numbers composed of 0 and 1, specifically including:
[0150] Taking the maximum and minimum values among all randomness data as a numerical interval, dividing the numerical interval into N parts to obtain N numerical segments, sorting the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; for each randomness data, if the value of the randomness data falls within the numerical range of a numerical segment S, then taking the sorting position of the numerical segment S as the feature transformation number of the randomness data; performing one-hot encoding on the feature transformation numbers of each randomness data to obtain encoded numbers composed of 0s and 1s.
[0151] This embodiment is the first feature transformation method, suitable for scenarios where the values of randomness data are continuously distributed. The following is an example to illustrate. Suppose the 1150 randomness data in the above example, their values are continuously distributed, the maximum value is 1, and the minimum value is 0; take N = 10, the numerical interval is 1000, divide the numerical interval 1000 into 10 equal segments, and the numerical ranges of each segment are [0, 0.1], [0.1, 0.2], [0.2, 0.3], [0.3, 0.4], [0.4, 0.5], [0.5, 0.6], [0.6, 0.7], [0.7, 0.8], [0.8, 0.9], [0.9, 1] respectively. Sorting by the numerical range size of the numerical interval, the corresponding sorting positions of the above numerical segments are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 respectively.
[0152] Match the 1150 randomness data with the above 10 numerical segments respectively. If the value of a randomness data a is 0.23 and falls within the [0.2, 0.3] segment, then taking the sorting position 3 of the [0.2, 0.3] segment as the feature transformation number of a; and so on. Perform segment matching on all randomness data respectively, then each randomness data will obtain a feature transformation number in the range of 1 to 10. Thus, the 1150 randomness data are divided into 10 parts, one part obtains the feature transformation number 1, one part obtains the feature transformation number 2, and so on, one part obtains the feature transformation number 10.
[0153] It should be noted that the sorting of the N numerical segments can be in ascending order or descending order, and the sorting method during malicious application detection only needs to be consistent with the sorting method during AI model training.
[0154] In addition, for the value at the boundary point of the numerical segment, it can be matched to the smaller numerical segment or the larger numerical segment according to the agreed rules, and the embodiments of the present invention do not limit this.
[0155] In addition, in the above example, the numerical interval 1000 is divided into 10 equal segments. If non-uniform segmentation is performed using specific rules, it also falls within the protection scope of the embodiments of the present invention.
[0156] One-hot encode the randomness data for which the feature transformation numbers have been obtained. For example, if the feature transformation number of the randomness data a is 1 and the total number of feature transformation numbers is 10, the one-hot encoding is 1000000000; if the feature transformation number of the value-type data b is 2, the one-hot encoding is 0100000000; and so on.
[0157] Based on any of the above optional embodiments, after respectively performing feature transformation on each randomness data, through one-hot encoding, it is converted into an encoded number composed of 0 and 1, specifically including:
[0158] Segment by the total number of all randomness data to obtain N numerical segments, and sort the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; sort each randomness data by size to obtain the sorting number of each randomness data; use the sorting position of the numerical segment corresponding to the sorting number of each randomness data as the feature transformation number of each randomness data; perform one-hot encoding on the feature transformation number of each randomness data, thereby obtaining an encoded number composed of 0 and 1.
[0159] This embodiment is the second feature transformation method, which is suitable for the scenario where the values of the randomness data are discretely distributed. The following is an example to illustrate. Assume the 1150 randomness data in the above example, whose values are discretely distributed. If the first feature transformation method is used, there may be no matching data in some numerical segments, and a large number of data may be matched in some numerical segments, which will have an adverse impact on the data operation of the AI model.
[0160] To avoid the above situation, the second feature transformation method uses the number of data for segmentation and matching. Assume the 1150 randomness data in the above example, take N = 10, then 10 numerical segments are obtained, and each numerical segment can match 115 randomness data; sort the 1150 randomness data by numerical size, then the 1st to 115th sorted randomness data match numerical segment 1 and obtain the feature transformation number 1, the 116th to 230th randomness data match numerical segment 2 and obtain the feature transformation number 2, and so on. The 1036th to 1150th randomness data match numerical segment 10 and obtain the feature transformation number 10. The one-hot encoding for the randomness data for which the feature transformation numbers have been obtained is the same as in the above embodiment and will not be elaborated here.
[0161] It should be noted that in the above example, the number of data 1150 is evenly divided into 10 segments. If non-uniform segmentation is performed using specific rules, it also falls within the protection scope of the embodiments of the present invention.
[0162] Based on any of the above optional embodiments, the feature transformation of each random degree data is respectively performed, and then converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including:
[0163] Part of all random degree data is segmented according to the number of data to obtain n numerical segments, where n is an integer greater than 0; some random degree data segmented according to the number of data in the random degree data is eliminated, and the numerical interval is divided into Nn parts with the maximum value and the minimum value in the remaining data as the numerical interval to obtain Nn numerical segments, where N is an integer greater than n; the N numerical segments are sorted according to the numerical size to obtain the ranking position of each numerical segment; each random degree data is sorted according to the size, and matched with the N numerical segments respectively, and the ranking position of the matched numerical segments is used as the characteristic transformation number of the random degree data; the characteristic transformation number of each random degree data is uniquely encoded to obtain a coding number composed of 0 and 1.
[0164] This embodiment is the third feature transformation method, which is suitable for scenarios where the values of randomness data have both discrete distribution and continuous distribution, and is explained below by example. It should be noted that the text description order of the third feature transformation method mentioned above is not used to limit the order of steps in actual implementation; in fact, all methods of segmenting by mixing numerical value size and data number, and matching by mixing numerical value size and data number, all fall within the protection scope of the embodiments of the present invention.
[0165] Assuming the 1150 random degree data in the above example, the 1150 random degree data are sorted by numerical value, wherein the first segment and the last segment are relatively discrete, and the data in the middle are relatively continuous; then the first segment, i.e., x% of the total data, is taken as the first numerical segment, and the sorting position is 1, and the last segment, i.e., y% of the total data, is taken as the Nth numerical segment, and the sorting position is N. Take N=10, x=10, y=20, 1150*10%=115, 1150*20%=230, then the first numerical segment can match 115 random degree data, and the Nth numerical segment can match 230 random degree data; after sorting the 1150 random degree data by numerical value, the 1st to 115th random degree data are matched to the first numerical segment, and the characteristic transformation number 1 is obtained, and the 921st to 1150th random degree data are matched to the Nth numerical segment, and the characteristic transformation number 10 is obtained.
[0166] Sort the 1150 randomness data, and take the maximum and minimum values among the 116th to 920th randomness data in the middle as the numerical interval. Divide the numerical interval into 10 - 2 segments, that is, 8 numerical segments, by an average or non - average method. The sorting positions are 2, 3, 4, 5, 6, 7, 8, 9 respectively. Match each randomness data among the 116th to 920th randomness data with the numerical segments 2 - 9. If the value of the randomness data is within the numerical range of a numerical segment S, then take the sorting position of the numerical segment S as its feature transformation number.
[0167] For the randomness data that has obtained the feature transformation number, the one - hot encoding is the same as the above - mentioned embodiment and will not be elaborated here.
[0168] It should be noted that in the above example, the first and last segments of the sorted randomness data are segmented according to the number of data, and the middle part is segmented according to the numerical range. It is also possible to segment the first and last segments according to the numerical range according to the needs of time application, and segment the middle part according to the number of data, and so on. As long as the data segmentation includes both segmentation according to the numerical range and segmentation according to the number of data, all fall within the protection scope of the embodiments of the present invention.
[0169] The embodiments of the present invention perform feature transformation on the extracted key features and then perform one - hot encoding, which is based on the data itself. During the training of the AI model, while ensuring that the information contained in the data can be maximally extracted, the final effect of the AI model is guaranteed; during the detection of malicious application programs, the accuracy of detection is maximally guaranteed.
[0170] Based on any of the above - mentioned optional embodiments, the network structure of the AI model in the embodiments of the present invention includes, but is not limited to, convolutional neural networks, recurrent neural networks, deep learning networks, machine learning networks, etc. It mainly includes four parts: an input layer, a factorization machine layer, a hidden layer, and an output layer. Figure 3 It is a schematic diagram of the network structure of the AI model described in the embodiments of the present invention. As Figure 3 shown, the network structure of the AI model includes four parts: an input layer (Input Layers), a factorization machine layer (Factorization Machines), a hidden layer (Hidden Layers), and an output layer (Output Layers);
[0171] The input layer is used to receive a digital feature vector and an application malicious label as input data; the decomposition machine layer is used to extract low-order features from the input data and perform calculations based on the low-order features; the hidden layer is used to extract high-order features from the input data, calculate the malicious features of the application according to the high-order features, and segment the malicious features of applications with different maliciousness from a high-dimensional space; the output layer is used to merge the calculation results of the decomposition machine layer and the calculation results of the hidden layer, and output the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application.
[0172] Figure 4 Schematic diagram of the AI model training process according to an embodiment of the present invention. As Figure 4 shown, the AI model according to the embodiment of the present invention is trained by the following method:
[0173] 400. Obtain an application installation package sample and the malicious label of the application installation package sample;
[0174] 401. Analyze the second target file in the application installation package sample, and extract the second key feature in the second target file; the form of the key feature is a string, and the second key feature includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;
[0175] 402. Obtain the word randomness of the second key feature, and splice the word randomness into a second digital feature vector;
[0176] 403. Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.
[0177] It should be noted that the second key feature and the second digital feature vector in the embodiment of the present invention are only for making a noun distinction from the key feature and the digital feature vector mentioned above, and the "second" has no actual meaning; the actual meanings of the second key feature and the key feature are exactly the same, and the actual meanings of the second digital feature vector and the digital feature vector are exactly the same.
[0178] In the embodiments of the present invention, the application program installation package sample data and malicious label data obtained in step 400 may be the results of manual analysis, where samples are collected manually and maliciousness is judged manually. In the embodiments of the present invention, a large number of complete Android platform application programs are collected, and the characteristics of malicious trojans in the application programs are analyzed manually, and then the rules for judging the maliciousness of the application programs are extracted. After that, based on the rules, the maliciousness of the collected Android platform application programs is judged to obtain the malicious labels corresponding to the application program installation package samples.
[0179] As described above, the processing procedures in the AI model training stage and the malicious application program detection stage in the embodiments of the present invention are exactly the same, including key feature extraction and calculation of the word randomness of the key features; if during AI model training, the randomness data of the key features is directly concatenated into a digital feature vector as the input data during AI model training, then in the malicious application program detection stage, the randomness data of the key features is also directly concatenated into a digital feature vector as the input data during AI model training; if during AI model training, the randomness data of the key features is subjected to feature transformation and one-hot encoding and then concatenated into a digital feature vector as the input data during AI model training, then in the malicious application program detection stage, the randomness data of the key features is also subjected to feature transformation and one-hot encoding and then concatenated into a digital feature vector as the input data during AI model training; the processing steps of all optional embodiments in the AI model training stage and the malicious application program detection stage are exactly the same, so as to ensure the detection effect of the AI model.
[0180] The processing method of step 401 of the AI model training method in the embodiments of the present invention is exactly the same as that of step 101 of the malicious application program detection method based on randomness and all optional embodiments of step 101; if there are multiple different optional embodiments, keeping the optional embodiments adopted in step 101 and step 401 consistent can achieve the effect of malicious application detection, which will not be elaborated here.
[0181] The processing method of step 402 of the AI model training method in the embodiments of the present invention is exactly the same as that of step 102 of the malicious application program detection method based on randomness and all optional embodiments of step 102; if there are multiple different optional embodiments, keeping the optional embodiments adopted in step 102 and step 402 consistent can achieve the effect of malicious application detection, which will not be elaborated here.
[0182] Step 403 of the AI model training method according to the embodiment of the present invention is different from step 103 of the malicious application detection method based on randomness in that the input data of step 403 includes not only the digital feature vector spliced by the randomness of key features, but also the malicious label of the application installation package sample. The AI model is trained by the digital feature vector and the malicious label of the application installation package sample, so that the AI model has the ability to identify malicious applications according to the digital feature vector.
[0183] The factorization machines layer of the AI model according to the embodiment of the present invention is used to extract low-order features in the data. The maliciousness of an APP is often a series of actions, and only by cooperating with each other in terms of behavior, permissions, content, etc. can the malicious behavior be completed. Therefore, the combination of features can often reflect the maliciousness of the APP to a large extent. The factorization machines mainly extract feature combinations through the inner product of latent variables for each dimension of features to ensure that the model can preserve the combination information between features.
[0184] Preferably, the hidden layers of the AI model according to the embodiment of the present invention is a feedforward neural network, including hidden units and sigmoid functions; there are 6 hidden layers in the hidden layer, and the number of neuron nodes in each layer ranges from 32 to 256. Specifically, the number of neuron nodes in each layer can be 64, 128, 256, 128, 64 or 32, which are used to extract high-order features in the data and divide the features of APPs with different maliciousness from a high-dimensional space.
[0185] The output layers of the AI model according to the embodiment of the present invention merge the forward calculation results of the factorization machines layer and the hidden layer, and the final number of output nodes is 2 or 1: if it is 2, they are the probability that this sample is a non-malicious application and the probability that it is a malicious application respectively; if it is 1, it is the probability that this sample is a non-malicious application or the probability that it is a malicious application.
[0186] Preferably, by setting the probability thresholds of malicious applications and non-malicious applications, the detection requirements of low false alarm scenarios can be met.
[0187] Finally, the model is trained by the method of backpropagation; taking cross-validation as a means and accuracy, recall rate, precision rate, and F value as criteria, considering the influence of time factors and business background, the parameters, number of iterations, model structure, etc. of the AI model are adjusted, and finally an AI model that meets the expected requirements is obtained.
[0188] Based on any of the above optional embodiments, in step 403, inputting the second digital feature vector and the malicious label of the application installation package sample into the established AI model for training the AI model specifically includes:
[0189] 403.1, converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector;
[0190] 403.2, inputting the second digital feature vector and the malicious label of the second digital feature vector into the established AI model for training the AI model.
[0191] In the embodiment of the present invention, the digital feature vector is composed of the randomness data of the key features extracted from the application installation package sample, or is composed of the randomness data after feature transformation. When training the AI model, if the digital feature vector is directly composed of the randomness data of the key features, then when detecting malicious application programs, the digital feature vector is also directly composed of the randomness data of the key features; if the digital feature vector is composed of the numbers after feature transformation and one-hot encoding of the randomness data when training the AI model, then when detecting malicious application programs, the digital feature vector is also composed of the numbers after feature transformation and one-hot encoding of the randomness data.
[0192] The malicious label of the application installation package sample is the result of tagging the maliciousness of the APP after analyzing a large number of mainstream malicious software on the Android platform. An application installation package sample corresponds to a digital feature vector and also corresponds to a malicious label. Therefore, the digital feature vector of the application installation package sample can be associated with the malicious label, and the malicious label of the application installation package sample can be converted into the malicious label of the digital feature vector, that is, one digital feature vector corresponds to one malicious label.
[0193] In this way, one input data of the AI model in the embodiment of the present invention is the digital feature vector, and the other input data is the malicious label corresponding to the digital feature vector, and the AI model is trained.
[0194] In the embodiment of the present invention, the digital feature vector is associated with the malicious label, and by tagging the digital feature vector, the generalization ability of the AI model can be improved.
[0195] Based on any of the above optional embodiments, in step 403.1, converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes:
[0196] Based on the malicious labels of all application installation package samples, use the malicious label of any one application installation package sample as the malicious label of the second digital feature vector corresponding to the any one application installation package sample;
[0197] Based on all the second digital feature vectors and their corresponding malicious labels, remove duplicates from the data where the second digital feature vectors and their corresponding malicious labels are exactly the same, and obtain the deduplicated second digital feature vectors and the malicious labels of the second digital feature vectors.
[0198] In the embodiment of the present invention, converting the malicious label of the application installation package sample into the malicious label of the digital feature vector will inevitably have data where the digital feature vector is exactly the same as the malicious label. For example, application A is malicious after analysis, application B is malicious after analysis, and application C is malicious after analysis; assume that after application A, B, and C extract key features and perform feature transformation, the obtained digital feature vectors are exactly the same, assume it is 0010000000, and the malicious labels of A, B, and C are all malicious, assume it is 1, then there will be three exactly the same data (0010000000, 1). In the embodiment of the present invention, removing duplicates from the data where the digital feature vector and its corresponding malicious label are exactly the same can greatly reduce the data volume and relieve the data processing pressure of the AI model. After laboratory scenario testing, after data deduplication, the data volume is reduced by about more than one order of magnitude.
[0199] Furthermore, if for multiple identical digital feature vectors, the corresponding malicious labels include malicious and non-malicious, and if the number of malicious ones is greater than the number of non-malicious ones, then mark this digital feature vector as malicious, otherwise mark it as non-malicious.
[0200] For example, for application A, B, C, D, and E, after extracting key features and performing feature transformation, the obtained digital feature vectors are exactly the same, assume it is 0010000000, and the malicious labels of A, B, and C are all malicious, assume it is 1, and the malicious labels of D and E are both non-malicious, assume it is 0, then the number of the digital feature vector 0010000000 corresponding to the malicious label 1 is greater than the number corresponding to the malicious label 0, so mark the digital feature vector 0010000000 as 1.
[0201] In this embodiment, according to the number of malicious and non-malicious corresponding to the malicious label of the digital feature vector, the digital feature vector is deduplicated and re-labeled, further reducing the data volume, relieving the data processing pressure of the AI model, and at the same time further improving the generalization ability of the AI model.
[0202] In summary, in the embodiments of the present invention, based on the analysis of Android platform applications, several key features such as application package names, program names, developer information, language information, behaviors, and permissions are extracted from three dimensions: behavior, permission, and content. Secondly, the randomness of the key features is calculated, and further data preprocessing such as feature transformation is performed on the randomness data. Thirdly, a model redesigned using the DeepFM algorithm in the deep learning method is used for multiple rounds of model training, and the evaluation strategy of the model effect is used to make the model effect reach the expected goal. Finally, the features of the application with unknown maliciousness are extracted and predicted through the model, and then the maliciousness determination of the application is completed.
[0203] Among them, the key features are the cornerstone of AI model training and the basis for malicious application detection. The key features include information such as specific behaviors, permissions, sensitive strings, and certificates of the application. Malicious applications have a strong distinguishability from normal applications in these specific dimensions, which is an important basis for judging the maliciousness of the application. A large number of keywords are obtained by verifying and parsing the application installation package, and the randomness algorithm is applied to the large number of keywords to obtain the randomness data of the keywords. All the randomness data is concatenated to obtain a vector, which is an important basis for judging whether the application is suspected of being batch-generated by malicious tools and whether these batch-generated applications carry viruses.
[0204] Furthermore, in the embodiments of the present invention, by performing feature transformation on the randomness data, from the perspective of the model, while ensuring that the information contained in the data can be maximally extracted, the final effect of the AI model is ensured.
[0205] Model design is the key point for the success of the AI model. Designing an appropriate model can maximize the conversion of the information in the data into knowledge, and then solidify the knowledge for the judgment of the maliciousness of the application.
[0206] In summary, the embodiments of the present invention perform static feature extraction and data statistical features on a large number of Android platform applications, redesign the algorithm based on the DeepFM algorithm in the deep learning method, and then perform model training on the data after feature extraction. Finally, an AI model for detecting malicious Android platform applications is obtained. The malicious detection of the application is performed through the AI model, which does not rely on the update of the malicious feature library and has strong learning ability, providing a new direction for the malicious detection of Android platform applications. The present invention solves the problems existing in the traditional malicious application detection based on manual extraction rules, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed. Moreover, the embodiments of the present invention have higher accuracy and timeliness.
[0207] Figure 5This is a malicious application detection device based on randomness according to an embodiment of the present invention. An embodiment of the present invention also provides a malicious application detection device based on randomness, as Figure 5 shown, including a key feature extraction module 501, a randomness acquisition module 502, and a malicious application detection module 503:
[0208] The key feature extraction module 501 parses the target file in the application installation package to be detected, extracts the key features in the target file, the form of the key features is a string, and the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;
[0209] The randomness acquisition module 502 acquires the word randomness of the key features and splices the word randomness into a digital feature vector;
[0210] The malicious application detection module 503 inputs the digital feature vector into a trained AI model to obtain a maliciousness detection result for the application to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non - malicious application.
[0211] The malicious application detection device based on randomness according to an embodiment of the present invention is used to execute Figure 1 the technical solution of the embodiment of the malicious application detection method based on randomness shown, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0212] Figure 6 This is a schematic diagram of the electronic device framework according to an embodiment of the present invention. Please refer to Figure 6, an embodiment of the present invention provides an electronic device, including: a processor 610, a communications interface 620, a memory 630, and a bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following method, including: parsing the target file in the application program installation package to be detected, extracting the key features in the target file, the form of the key features is a string, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension; obtaining the word randomness of the key features, and splicing the word randomness into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.
[0213] An embodiment of the present invention discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments, for example, including: parsing the target file in the application program installation package to be detected, extracting the key features in the target file, the form of the key features is a string, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension; obtaining the word randomness of the key features, and splicing the word randomness into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.
[0214] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: parsing a target file in a to-be-detected application installation package, extracting key features in the target file, the form of the key features being strings, and the key features including at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; obtaining the word randomness of the key features, and splicing the word randomness into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the to-be-detected application; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.
[0215] Those of ordinary skill in the art can understand that implementing the above device embodiment or method embodiment is merely illustrative. The processor and the memory may or may not be physically separated components, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0216] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a USB flash drive, a mobile hard disk, a ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting malicious application programs based on randomness, characterized in that, it includes: Analyze the target files in the installation package of the application program to be detected, and extract the key features in the target files. The form of the key features is a string, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension; Obtain the word randomness of the key features, and splice the word randomness into a digital feature vector; Input the digital feature vector into the trained AI model to obtain the maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application; The behavior dimension includes the behavior information when the application program runs; The permission dimension includes the permission information required by the application program when performing specific behaviors; The content dimension includes at least one of the following information: application package name, program name, developer information, and language information; The obtaining the word randomness of the key features and splicing the word randomness into a digital feature vector specifically includes: According to the alphabetical order of the key feature strings, sequentially obtain the adjacent letter frequencies of any two adjacent letters of the key feature strings; Based on the adjacent letter frequencies of any two adjacent letters, obtain the word randomness of the key features; Splice the word randomness into a digital feature vector; Among them, the adjacent letter frequency is to regard the string as a word of a preset language text, and obtain the adjacent frequency between two letters through the word randomness calculation rule of the preset language text; The word randomness is obtained through the adjacent letter frequency; The splicing the word randomness into a digital feature vector specifically includes: Based on each randomness data in the English word randomness and / or the Chinese pinyin randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, convert it into an encoded number composed of 0 and 1; Splice all the encoded numbers in a preset order into a digital feature vector composed of 0 and 1.
2. The method according to claim 1, characterized in that, The specifically performing feature transformation on each randomness data respectively, and then through one-hot encoding, converting it into an encoded number composed of 0 and 1 includes: Based on all the randomness data, obtain N numerical segments according to the first preset rule, and sort the N numerical segments according to the numerical size to obtain the sorting positions of each numerical segment. N is an integer greater than 0; Match each randomness data with the N numerical segments according to the second preset rule, so that each randomness data matches a numerical segment, and use the sorting position of the matched numerical segment as the feature transformation number of each randomness data; Perform one-hot encoding on the feature transformation numbers of each randomness data, so as to obtain an encoded number composed of 0 and 1.
3. The method according to claim 1, characterized in that, After respectively performing feature transformation on each randomness data, through one-hot encoding, it is converted into encoded numbers composed of 0 and 1, specifically including: Taking the maximum value and the minimum value among all randomness data as the numerical interval, dividing the numerical interval into N parts, obtaining N numerical segments, and sorting the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; For each randomness data, if the value of the randomness data is within the numerical range of a numerical segment S, then taking the sorting position of the numerical segment S as the feature transformation number of the randomness data; Performing one-hot encoding on the feature transformation numbers of each randomness data, thereby obtaining encoded numbers composed of 0 and 1.
4. The method according to claim 1, wherein, After respectively performing feature transformation on each randomness data, through one-hot encoding, it is converted into encoded numbers composed of 0 and 1, specifically including: Segmenting according to the total number of all randomness data to obtain N numerical segments, and sorting the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; Sorting each randomness data by size to obtain the sorting number of each randomness data; Taking the sorting position of the numerical segment corresponding to the sorting number of each randomness data as the feature transformation number of each randomness data; Performing one-hot encoding on the feature transformation numbers of each randomness data, thereby obtaining encoded numbers composed of 0 and 1.
5. The method according to claim 1, wherein, After respectively performing feature transformation on each randomness data, through one-hot encoding, it is converted into encoded numbers composed of 0 and 1, specifically including: Segmenting part of the data of all randomness data according to the number of data to obtain n numerical segments, where n is an integer greater than 0; Eliminating part of the randomness data segmented according to the number of data in the randomness data, taking the maximum value and the minimum value of the remaining data as the numerical interval, dividing the numerical interval into N - n parts to obtain N - n numerical segments, where N is an integer greater than n; Sorting the N numerical segments by numerical size to obtain the sorting positions of each numerical segment; Sorting each randomness data by size, respectively matching with the N numerical segments, and taking the sorting position of the matched numerical segment as the feature transformation number of the randomness data; Performing one-hot encoding on the feature transformation numbers of each randomness data, thereby obtaining encoded numbers composed of 0 and 1.
6. The method according to any one of claims 1 - 5, wherein, The network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method: Obtaining application program installation package samples and the malicious labels of the application program installation package samples; Analyze the second target file in the application installation package sample, and extract the second key feature in the second target file. The form of the key feature is a string, and the key feature includes at least one dimension information among the behavior dimension, permission dimension, and content dimension; Obtain the word randomness of the second key feature, and splice the word randomness into a second digital feature vector; Input the second digital feature vector and the malicious label of the application installation package sample into the established AI model, and train the AI model to obtain an AI model that meets the expected requirements.
7. The method according to claim 6, wherein, The step of inputting the second digital feature vector and the malicious label of the application installation package sample into the established AI model and training the AI model specifically includes: Convert the malicious label of the application installation package sample into the malicious label of the second digital feature vector; Input the second digital feature vector and the malicious label of the second digital feature vector into the established AI model, and train the AI model.
8. The method according to claim 7, wherein, The step of converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes: Based on the malicious labels of all application installation package samples, use the malicious label of any application installation package sample as the malicious label of the second digital feature vector corresponding to the any application installation package sample; Based on all the second digital feature vectors and their corresponding malicious labels, remove duplicates from the data with exactly the same second digital feature vectors and their corresponding malicious labels to obtain the deduplicated second digital feature vectors and the malicious labels of the second digital feature vectors.
9. An electronic device, wherein, comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 8 by invoking the program instructions.
10. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Word vector- and depth neural network-based Android malicious code detection method
CN108959924A
Malicious program detection method and device and related application
CN109726554A