A Feature-Based Malicious Application Detection Method and Device

By extracting the attack intentions and means of malware on Android platform and combining deep learning algorithms, feature-based malicious application detection is realized, solving the problem of insufficient malicious detection capabilities in the existing technology, and improving the accuracy and timeliness of detection.

CN113971283BActive Publication Date: 2025-06-10WUHAN ANTIY MOBILE SECURITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010722234.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-24
Publication Date
2025-06-10
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

Existing mobile security software relies on the update of malicious feature library in malicious application detection, and learns that new malicious detection capabilities are weak, resulting in the inability to effectively protect new malicious applications.

Method used

A feature-based malicious application detection method is adopted. By analyzing the target files in the application installation package, key features of behavior dimensions, permission dimensions and content dimensions are extracted, word randomness and numerical data are calculated, and a deep learning algorithm is used to train an AI model for malicious detection.

Benefits of technology

It improves the accuracy and timeliness of malicious application detection, solves the problems of difficulty in extracting rules, low coverage, poor scalability, and easy to be bypassed in traditional methods, and does not rely on the update of the malicious feature library.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971283B_ABST
    Figure CN113971283B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method and device for detecting malicious application programs based on features. The method includes parsing a target file in a to-be-detected application program installation package, extracting key features in the target file, where the form of the key features is a string; obtaining the word randomness of some key features, obtaining numerical data of the remaining key features, and splicing the word randomness and the numerical data into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result for the to-be-detected application program; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application. It solves the problems existing in traditional malicious application program detection, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed, and has higher accuracy and timeliness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of mobile network security, and in particular, to a method and device for detecting malicious application programs based on features. Background Art

[0002] At present, many security vendors have also entered the mobile security field. The basic principle of these software for antivirus is to confirm intrusion behaviors by matching known malicious trojan features and perform active defense in ways such as firewalls and dynamic monitoring. However, the disadvantage is that it depends on the update of the malicious feature library and has a weak ability to learn new malicious detections.

[0003] However, new malicious application programs emerge in an endless stream, and their malicious natures are different. Relying on the malicious feature library for malicious detection, it is difficult to achieve an ideal security protection effect due to the untimely update of the malicious feature library. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, embodiments of the present invention provide a method and device for detecting malicious application programs based on features. On the basis of analyzing a large number of mainstream malicious software on the Android platform, the attack intentions and means of malicious software on the Android platform are summarized, and AI detection of the maliciousness of application programs is realized through a deep learning algorithm, providing a new direction for the malicious detection of application programs on the Android platform.

[0005] In a first aspect, embodiments of the present invention provide a method for detecting malicious application programs based on features, including:

[0006] Analyze the target files in the installation package of the application program to be detected, and extract the key features in the target files. The key features include at least one type of dimension information among the behavior dimension, permission dimension, and content dimension;

[0007] Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a digital feature vector;

[0008] Input the digital feature vector into a trained AI model to obtain a maliciousness detection result for the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0009] Further, the obtaining of the word randomness of the some key features specifically includes:

[0010] According to the alphabetical order of the partial key feature strings, sequentially obtain the letter adjacent frequencies of any two adjacent letters of the partial key feature strings;

[0011] Obtain the randomness of words of the key feature based on the adjacent letter frequency of any two adjacent letters among them.

[0012] Wherein, the adjacent letter frequency is obtained by taking the string as a word of a preset language text and calculating the adjacent frequency between two letters through the word randomness calculation rule of the preset language text; the word randomness is obtained through the adjacent letter frequency.

[0013] Furthermore, the network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method:

[0014] Obtain the application program installation package samples and the malicious labels of the application program installation package samples.

[0015] Parse the second target file in the application program installation package sample, and extract the second key feature in the second target file, where the key feature includes at least one dimension information among the behavior dimension, the permission dimension, and the content dimension.

[0016] Obtain the randomness of words of some key features, obtain the numerical data of the remaining key features, and splice the randomness of words and the numerical data into a digital feature vector.

[0017] Input the second digital feature vector and the malicious label of the application program installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.

[0018] Furthermore, the step of inputting the second digital feature vector and the malicious label of the application program installation package sample into the built AI model and training the AI model specifically includes:

[0019] Convert the malicious label of the application program installation package sample into the malicious label of the second digital feature vector.

[0020] Input the second digital feature vector and the malicious label of the second digital feature vector into the built AI model, and train the AI model.

[0021] In a second aspect, an embodiment of the present invention provides an electronic device, including:

[0022] At least one processor; and

[0023] At least one memory communicatively connected to the processor, wherein:

[0024] The memory stores program instructions executable by the processor, and the processor can execute the method for detecting malicious application programs based on features described in the first aspect of the embodiments of the present invention and the methods described in any optional embodiments thereof by invoking the program instructions.

[0025] In a third aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions execute the method for detecting malicious application programs based on features described in the first aspect of the embodiments of the present invention and the methods described in any optional embodiments thereof.

[0026] The method for detecting malicious application programs based on features provided by the embodiments of the present invention extracts key features of target files in the installation package of the application to be detected, obtains the randomness of words for some key features, obtains numerical data for the remaining key features, and splices all the obtained randomness data and numerical data into a digital feature vector; based on the already trained AI model, by performing operations on the digital feature vector through the AI model, the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application can be obtained. The embodiments of the present invention solve the problems existing in the traditional detection of malicious application programs based on manually extracted rules, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed, and have higher accuracy and timeliness for malicious program detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a schematic flow chart of the method for detecting malicious application programs based on features described in the embodiments of the present invention;

[0029] Figure 2 It is a schematic diagram of numerical segmentation in the embodiments of the present invention;

[0030] Figure 3 It is a schematic diagram of the network structure of the AI model described in the embodiments of the present invention;

[0031] Figure 4 It is a schematic flow chart of the training of the AI model in the embodiments of the present invention;

[0032] Figure 5 It is a device for detecting malicious application programs based on features in the embodiments of the present invention;

[0033] Figure 6Schematic diagram of the framework of the electronic device according to the embodiment of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] In view of the problems in the prior art, the embodiment of the present invention redesigns the algorithm based on the DeepFM algorithm in the deep learning method to obtain an artificial intelligence (AI) model; based on the extraction of key features from a large number of Android platform application programs, the AI model is trained by the extracted key features, and finally an AI model for detecting malicious application programs on the Android platform is obtained, where the training process includes: extracting key features from sample data (sample of application program installation packages); performing feature transformation on the extracted key features to obtain a digital feature vector corresponding to the sample data; using the digital feature vector corresponding to the sample data and the malicious label corresponding to the sample data as input data of the AI model to train the AI model, and evaluating the effect through a model evaluation strategy to obtain an AI model meeting the expected requirements.

[0036] When performing malicious detection on an application program, key features are extracted from the target file in the application program installation package to be detected, the word randomness is calculated for some key features, and numerical data (further feature transformation is performed) is obtained for the remaining key features, and they are concatenated into a digital feature vector corresponding to the application program installation package to be detected; the digital feature vector corresponding to the application program installation package to be detected is used as input data of the AI model, and the AI model outputs the probability that the application program installation package to be detected is a malicious application and / or the probability that it is a non-malicious application after calculation.

[0037] In the embodiment of the present invention, in both the AI model training stage and the malicious application program detection stage, the word randomness is calculated for some key features, and numerical data (further feature transformation is performed) is obtained for the remaining key features, and their processing methods are exactly the same; in all optional solutions of the embodiment of the present invention, the solution adopted in the malicious detection stage of the application program is consistent with the solution adopted in the AI model training stage to obtain the expected detection effect.

[0038] Next, from the perspective of the malicious application program detection stage, the feature-based malicious application program detection method described in the embodiment of the present invention will be described in detail.

[0039] Figure 1 This is a schematic flowchart of the feature - based malicious application detection method according to an embodiment of the present invention. As Figure 1 shown, the feature - based malicious application detection method includes:

[0040] 101. Analyze the target file in the installation package of the application to be detected, and extract the key features in the target file. The key features include at least one type of dimension information among the behavior dimension, permission dimension, and content dimension;

[0041] 102. Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a digital feature vector;

[0042] 103. Input the digital feature vector into the trained AI model to obtain the maliciousness detection result of the application to be detected. The AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non - malicious application.

[0043] When the embodiment of the present invention performs malicious detection on an application, first, obtain the installation package of the application to be detected, analyze the target file in the installation package of the application to be detected, and extract the key features in the target file. The key features include at least one type of dimension information among the behavior dimension, permission dimension, and content dimension. Each type of dimension information may include one or more different pieces of information. For example, the behavior dimension may include Behavior 1, Behavior 2, Behavior 3, etc., the permission dimension may include Permission 1, Permission 2, Permission 3, Permission 4, etc., and the content dimension may include Content 1, Content 2, Content 3, etc. The forms of all the obtained key features are strings.

[0044] Specifically, the behavior dimension described in the embodiment of the present invention includes the behavior information when the application is running; the permission dimension includes the permission information required when the application performs a specific behavior; the content dimension includes at least one of the following information: application package name, program name, developer information, language information, total file size of the application, number of files included in the application, size of a specific file in the application, number of specific components in the application, number of metadata, number of resource strings, number of supported languages, and application hardening information.

[0045] Step 102 divides the extracted key features into two parts, calculates the word randomness for some key features, and obtains the numerical data for the remaining key features. The following is a description of each part respectively.

[0046] I. Calculate the word randomness for some key features:

[0047] Implementing the word randomness in the embodiments of the present invention may include the word randomness of one or more languages and characters, such as English word randomness, Japanese word randomness, French word randomness, Chinese Pinyin randomness, and so on. When applied, one type of word randomness can be used according to the application environment, or multiple types of word randomness can be used simultaneously.

[0048] Taking the English word randomness and the Chinese Pinyin randomness as examples, for one or more of the key features in some key features, the English word randomness of each key feature can be calculated, and all the English word randomness can be concatenated into a digital feature vector; or the Chinese Pinyin randomness of each key feature can be calculated, and all the Chinese Pinyin randomness can be concatenated into a digital feature vector; or the English word randomness and the Chinese Pinyin randomness of each key feature can be calculated simultaneously, and all the English word randomness and the Chinese Pinyin randomness can be concatenated into a digital feature vector.

[0049] In the analysis of a large number of application program samples, it is found that many malicious samples containing viruses are automatically generated based on special tools. Since such samples are automatically generated, the keywords in the form of strings such as the package name, program name, and developer information of these samples often have strong randomness. Therefore, analyzing the randomness of application program keywords is a very effective virus detection method. Randomness is a determination of whether a word is a reasonable word or a combination of random letters. The closer a word is to a common word, the lower its randomness, and vice versa. The higher the randomness, the greater the possibility of a malicious application.

[0050] Second, obtain numerical data for the remaining key features:

[0051] Preferably, the application program package name, program name, developer information, and language information in the content dimension of the embodiments of the present invention can be used as some key features for calculating the word randomness;

[0052] For other features in the content dimension and features in the behavior dimension and permission dimension, their numerical values can be obtained, such as the number of specific components in the application program, the size of specific files in the application program, the number of metadata, etc.; in the behavior dimension, if the corresponding behavior is obtained, this key feature can be converted into the number 1, and if the corresponding behavior is not obtained, this key feature can be converted into the number 0; similarly, in the permission dimension, if the corresponding permission is obtained, this key feature can be converted into the number 1, and if the corresponding permission is not obtained, this key feature can be converted into the number 0; thus, all the remaining key features can be converted into numerical data.

[0053] Based on the above processing, randomness data and numerical data are obtained, and all the randomness data and numerical data are directly concatenated into a digital feature vector. Embodiments of the present invention do not limit the specific concatenation order, as long as the concatenation order in the malicious application detection phase is consistent with that in the AI model training phase.

[0054] Embodiments of the present invention directly form a digital feature vector from the randomness data and numerical data of the key features, and use it as the input data of the AI model to judge malicious applications through the randomness of the key features of the application.

[0055] Step 103 inputs the digital feature vector into a pre-trained AI model, and through calculation, the maliciousness detection result of the application to be detected can be obtained. The maliciousness detection result described in embodiments of the present invention refers to the probability that the application to be detected is a malicious application and / or the probability that it is a non-malicious application. The AI model in embodiments of the present invention can output the probability that the application is a malicious application, or output the probability that the application is a non-malicious application, or output the probability that the application is a malicious application and the probability that it is a non-malicious application at the same time.

[0056] Assume that the AI model outputs the probability of a malicious application and the probability of a non-malicious application at the same time. If the former is greater than the latter, the application is non-malicious; otherwise, the application is malicious. For example, it outputs "[0.98654, 0.01346]", where 0.98654 is the probability of a non-malicious application and 0.01346 is the probability of a malicious application, indicating that this program is non-malicious, thus completing the judgment of the maliciousness of the application.

[0057] The method for detecting malicious applications based on features described in embodiments of the present invention extracts the key features of the target file in the installation package of the application to be detected, obtains the word randomness of some key features, obtains numerical data for the remaining key features, and concatenates all the obtained randomness data and numerical data into a digital feature vector; based on the already trained AI model, by operating on the digital feature vector through the AI model, the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application can be obtained. Embodiments of the present invention solve the problems existing in the traditional detection of malicious applications based on manually extracted rules, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed, and have higher accuracy and timeliness for malicious program detection.

[0058] Based on the above embodiments, step 101 of parsing the target file in the installation package of the application to be detected and extracting the key features in the target file specifically includes:

[0059] 101.1, parse the information of each sub-file in the installation package of the application to be detected to obtain the target file;

[0060] 101.2, in the target file, extract the target string that matches a specific keyword, which is the key feature; the target string contains at least one dimension information of the behavior dimension, permission dimension, and content dimension;

[0061] In step 101.1 of the embodiment of the present invention, the target file is obtained from the sub-files of the application installation package to be detected, and the target file can be files such as AndroidManifest.xml and classes.dex in the sub-files.

[0062] In the embodiment of the present invention, the key feature in the target file is extracted by keyword matching. The specific keyword in step 101.2 is a keyword related to behavior, permission, and content, and the target string matched by the specific keyword is the string information related to behavior, permission, and content. By matching all the specific keywords set in the embodiment of the present invention, all the information of the behavior dimension, permission dimension, and content dimension may be matched, or only part of the information may be matched.

[0063] The key feature in the embodiment of the present invention contains at least one dimension information of the behavior dimension, permission dimension, and content dimension; specifically, the extraction methods of each dimension information include but are not limited to the following descriptions, and each dimension information includes at least one of the descriptions:

[0064] In the behavior dimension, the behavior of the application refers to behaviors such as accessing local files, sending text messages, and connecting to the network when the program is running. Specifically, these behavior information exist in the classes.dex file in the form of specific class names, method names, and variable names. The behavior of the application is an important basis for virus detection. Many virus-infected applications will have several high-risk behaviors. Analyzing the list of high-risk behaviors of the application is an important basis for judging whether the application carries a virus. Based on the analysis of a large number of existing applications, the connection between the above key symbols and the application behavior can be found, and the specific combination of the key symbols and the application behavior can be corresponding to form a behavior feature library.

[0065] For example:

[0066] (1) The behavior "SMS_Send" indicates whether there is a behavior of sending text messages, and the corresponding behavior rule is to include one of the following "class -> method":

[0067] android / telephony / SmsManager->sendMultipartTextMessage;

[0068] android / telephony / SmsManager->sendTextMessage;

[0069] android / telephony / gsm / SmsManager->sendTextMessage;

[0070] android / telephony / gsm / SmsManager->sendMultipartTextMessage;

[0071] android / telephony / SmsManager->sendDataMessage;

[0072] android / telephony / gsm / SmsManager->sendDataMessage;

[0073] (2) The behavior "SMS_Scan" represents the behavior of SMS search, and the corresponding behavior rule is the combination of "class -> method" that contains:

[0074] android / database / Cursor->moveToFirst, and the static string contains the url: content: / / sms;

[0075] Or it contains the combination of "class -> method":

[0076] android / database / Cursor->moveToPosition, and the static string contains the url: content: / / sms;

[0077] In the dimension of permissions, when an application performs the behavior of reading files, it needs to declare the permission to read files first. Some of these permissions are sensitive. Therefore, whether certain permissions are declared or not has certain reference value for judging the maliciousness of the application. Preferably, the embodiments of the present invention obtain information about permission declarations from the "<uses-permission name = "xxx">" field of the AndroidManifest.xml file.

[0078] In the content dimension, the embodiments of the present invention analyze the AndroidManifest.xml file, which is a standard XML file after decryption. Query the <manifest package="xxx"> field, where xxx is the package name of the application; query the <application lable="xxx"> field, where xxx represents the index of the application name in the resources.arsc file. The application name can be obtained by parsing the resources.arsc through this index. The certificate file is located in the META-INF directory and is in the standard X509 format. The openssl-related interfaces can be called to parse the certificate and obtain the developer information of the application. OpenSSL is a large encryption and decryption library for common standard formats. Since the certificate belongs to the X509 standard format, the developer information can be obtained directly by calling the relevant interfaces for parsing.

[0079] Other characteristics in the content dimension may also include: the installation package size, obtained by parsing the installation package header; the size of the classes.dex file, obtained by extracting the classes.dex file; the number of sub-files in the installation package, obtained by parsing the installation package header; the number of Activity components, obtained by parsing the AndroidManifest.xml; the number of Service components, obtained by parsing the AndroidManifest.xml; the number of Receiver components, obtained by parsing the AndroidManifest.xml; the number of Meta-data, obtained by parsing the AndroidManifest.xml; the number of supported languages, obtained by parsing the resources.arsc; the number of resource strings, obtained by parsing the resources.arsc; and so on.

[0080] Based on the analysis of a large number of existing application programs, the embodiments of the present invention extract information in dimensions such as the specific behaviors, permissions, sensitive strings, and certificates of the application program included in the application installation package. Malicious application programs have a strong distinction from normal application programs in these dimensions, which is an important basis for judging the maliciousness of application programs.

[0081] Based on any of the above optional embodiments, obtaining the word randomness of the partial key features in step 102 specifically includes:

[0082] 102.1, according to the alphabetical order of the partial key feature strings, sequentially obtain the letter adjacent frequencies of any two adjacent letters of the partial key feature strings;

[0083] 102.2, based on the letter adjacent frequencies of any two adjacent letters, obtain the word randomness of the key features;

[0084] Among them, the adjacent letter frequency is obtained by taking the string as a word in a preset language text and calculating the adjacent frequency between two letters according to the word randomness calculation rule of the preset language text; the word randomness is obtained from the adjacent letter frequency.

[0085] For example, if the preset language is English, the adjacent letter frequency is obtained by taking the string as English and calculating the adjacent frequency between two letters according to the English word randomness calculation rule;

[0086] if the preset language is Chinese, the adjacent letter frequency is obtained by taking the string as Chinese pinyin and calculating the adjacent frequency between two letters according to the Chinese pinyin randomness calculation rule;

[0087] Correspondingly, the English word randomness is obtained from the adjacent letter frequency of the English word; the Chinese pinyin randomness is obtained from the adjacent letter frequency of the Chinese pinyin.

[0088] In the embodiments of the present invention, the word randomness calculation rule is obtained by analyzing the letter order of a large number of natural languages. The English word randomness calculation rule is obtained by analyzing the letter order of English words in a large number of natural languages. The Chinese pinyin randomness calculation rule is obtained by analyzing the letter order of Chinese pinyin in a large number of natural languages.

[0089] For an application program, the string randomness calculation can use any set of strings parsed from the installation package. By calculating the randomness of the string, it can be determined whether a string is a reasonable word or a random combination of letters, so as to determine whether the corresponding application program is malicious or non-malicious.

[0090] In practical applications, there are many randomness calculation rules, and the string rules in different application fields also have specific meanings. Therefore, the randomness calculation rules in different application fields are also different, and the embodiments of the present invention do not limit this. The key point of the embodiments of the present invention is that no matter what rule is used to calculate the randomness of the strings in this field, the lower the randomness, the greater the possibility that the string is a reasonable word; the higher the randomness, the greater the possibility that the string is a random combination of letters, and this is used as the basis for determining whether the application program is malicious or non-malicious.

[0091] By analyzing the alphabetical order of a string in different languages, the adjacent letter frequencies of any two letters can be obtained, and these adjacent letter frequencies can be pre - formed into a frequency library file. In the embodiments of the present invention, according to the alphabetical order of the key feature string to be calculated, the adjacent letter frequencies of adjacent letters are queried in turn, and then by synthesizing the occurrence frequencies of all adjacent letters, the overall adjacent frequency value of a string is obtained, named "entropy". The lower the entropy of a string, the higher the randomness of the string, indicating that the string is less like a regular word.

[0092] The calculation method of the entropy of the string in the embodiments of the present invention is as follows:

[0093] (1) Let entropy represent entropy, and set the initial entropy to 0;

[0094] (2) Traverse the string. For every two adjacent letters in the string, query the adjacent frequency of this two - letter combination in the frequency library file, and record it as prob;

[0095] (3) Update entropy according to the following formula:

[0096] entropy = entropy - prob * log 2 prob;

[0097] (4) Repeat steps (2) and (3) until the string traversal ends.

[0098] The following uses English words and Chinese pinyin as examples to illustrate randomness.

[0099] Suppose the package name of application 1 is: com.iizirgui.vcpnocuc;

[0100] The developer information is:

[0101] CN = hmqmawavzg.oaa, OU = tdbkusvGGE, O = ppCIFivngI,

[0102] L = fRitORayrU, ST = DtTJWEpbqs, C = QL,

[0103] Regarding the package name com.iizirgui.vcpnocuc as English, obtaining the adjacent frequency between every two letters according to the English word randomness calculation rule, and obtaining the English entropy of its package name as 0.160173 according to the calculation method of the string entropy; regarding the package name com.iizirgui.vcpnocuc as Chinese pinyin, obtaining the adjacent frequency between every two letters according to the Chinese pinyin randomness calculation rule, and obtaining the Chinese pinyin entropy as 0.118652 according to the calculation method of the string entropy;

[0104] Similarly, the English entropy of the CN segment of the developer information is 0.117621, and the Chinese pinyin entropy is 0.093654; the English entropy of the OU segment of the developer information is 0.130641, and the Chinese pinyin entropy is 0.113263; the English entropy of the O segment of the developer information is 0.222635, and the Chinese pinyin entropy is 0.099876.

[0105] As mentioned above, the randomness of English words and the randomness of Chinese pinyin are obtained simultaneously, and the randomness of English words and the randomness of Chinese pinyin are spliced into a digital feature vector in a preset order; as described above, the embodiment of the present invention does not limit the splicing order, and one of the splicings is as follows:

[0106] [0.160173, 0.118652, 0.117621, 0.093654, 0.130641, 0.113263, 0.222635, 0.099876].

[0107] As mentioned above, the partial key features mainly include at least one of the application package name, program name, developer information, and language information in the content dimension of the embodiment of the present invention.

[0108] Preferably, the embodiment of the present invention obtains the randomness of English words and / or the randomness of Chinese pinyin of the partial key features, that is: obtains the randomness of English words of the partial key features, obtains the randomness of Chinese pinyin of the partial key features, and simultaneously obtains the randomness of English words and the randomness of Chinese pinyin of the partial key features.

[0109] Specifically, when obtaining the randomness of English words, the key feature string is used as an English word, and the letter adjacent frequency of any two adjacent letters is obtained in turn, and then according to the letter adjacent frequency of all letters and the calculation method of entropy, the randomness of English words of the key feature string is obtained.

[0110] When obtaining the randomness of Chinese pinyin, the key feature string is used as Chinese pinyin, and the letter adjacent frequency of any two adjacent letters is obtained in turn, and then according to the letter adjacent frequency of all letters and the calculation method of entropy, the randomness of Chinese pinyin of the key feature string is obtained.

[0111] Based on any of the above optional embodiments, step 102 of obtaining the numerical data of the remaining partial key features specifically includes:

[0112] 102.3, based on the key features in the content dimension of the remaining partial key features, convert the extracted key features into first numerical data, and the first numerical data is the value represented by the key features;

[0113] 102.4. Based on the key features of the behavior dimension and the permission dimension in the remaining key features, convert the extracted key features into second numerical data, where the second numerical data is the number 0 or 1.

[0114] As mentioned above, in the content dimension, the application package name, program name, developer information, and language information can be used as some of the key features for calculating the randomness of words; for other features in the content dimension and the features in the behavior dimension and permission dimension, their numerical values can be obtained as the first numerical data, such as the number of specific components in the application, the size of specific files in the application, the number of metadata, etc.; in the behavior dimension, if the corresponding behavior is obtained, this key feature can be converted into the number 1, and if the corresponding behavior is not obtained, this key feature can be converted into the number 0; similarly, in the permission dimension, if the corresponding permission is obtained, this key feature can be converted into the number 1, and if the corresponding permission is not obtained, this key feature can be converted into the number 0. After the key features in the behavior dimension and permission dimension are converted, the second numerical data is obtained; thus, all the remaining key features can be converted into numerical data.

[0115] It should be noted that the embodiments of the present invention do not limit the execution order of steps 102.1, 102.2, 102.3, and 102.4.

[0116] Based on any of the above optional embodiments, step 102 of splicing the word randomness and the numerical data into a digital feature vector specifically includes:

[0117] 102.5. Based on each data in the word randomness and the first numerical data, perform feature transformation on each data respectively, and then through one-hot encoding, convert it into an encoded number composed of 0 and 1.

[0118] 102.6. Splice the encoded numbers after converting all the randomness data and the first numerical data, and the second numerical data into a digital feature vector in a preset order.

[0119] Step 102.5 of the embodiments of the present invention includes two parts:

[0120] The first part: Perform feature transformation on the first numerical data, and then perform one-hot encoding to obtain the encoded number of each first numerical data.

[0121] The second part: According to the word randomness of some key feature strings, perform feature transformation, and then perform one-hot encoding to obtain the encoded number of each randomness data.

[0122] Preferably, the embodiments of the present invention obtain the English word randomness and / or Chinese pinyin randomness of some key features, including three methods:

[0123] (1) Based on each randomness data in the English word randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, it is converted into encoded numbers consisting of 0 and 1;

[0124] (2) Based on each randomness data in the Chinese pinyin randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, it is converted into encoded numbers consisting of 0 and 1;

[0125] (3) Based on each randomness data in the English word randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, it is converted into encoded numbers consisting of 0 and 1; based on each randomness data in the Chinese pinyin randomness, after performing feature transformation on each randomness data respectively, through one-hot encoding, it is converted into encoded numbers consisting of 0 and 1.

[0126] (1) and (2) perform feature transformation on the English word randomness and the Chinese pinyin randomness respectively, and (3) performs feature transformation on the English word randomness and the Chinese pinyin randomness simultaneously. Each randomness data will obtain a feature transformation number, and then the feature transformation numbers are one-hot encoded, and finally spliced into a digital feature vector consisting of 0 and 1. Through feature transformation, while retaining more feature information, the generalization ability of the AI model can be improved.

[0127] The encoded numbers after the conversion of the randomness data and the first numerical data in the embodiments of the present invention consist of 0 and 1, and the second numerical data is 0 or 1. Therefore, the spliced digital feature vectors are all composed of 0 and 1. Corresponding to the above three methods, step 102.6 in the embodiments of the present invention splicing all the encoded numbers and the second numerical data into a digital feature vector according to a preset order also includes three types:

[0128] (1) Splice the encoded numbers after the conversion of the English word randomness, the encoded numbers after the conversion of the first numerical data, and the second numerical data into a digital feature vector according to a preset order;

[0129] (2) Splice the encoded numbers after the conversion of the Chinese pinyin randomness, the encoded numbers after the conversion of the first numerical data, and the second numerical data into a digital feature vector according to a preset order;

[0130] (3) Splice the encoded numbers after the conversion of the English word randomness, the encoded numbers after the conversion of the Chinese pinyin randomness, the encoded numbers after the conversion of the first numerical data, and the second numerical data into a digital feature vector according to a preset order.

[0131] Similarly, the embodiments of the present invention do not limit the splicing order of the digital feature vectors after feature transformation and one-hot encoding, as long as the splicing order in the malicious application detection phase is consistent with that in the AI model training phase.

[0132] One-hot encoding, also known as one-hot effective encoding, is a method of using an N-bit status register to encode N states. Each state has its own independent register bit, and at any time, only one of them is effective.

[0133] For example, encoding six states:

[0134] The natural sequence code is: 000, 001, 010, 011, 100, 101;

[0135] The one-hot encoding is: 000001, 000010, 000100, 001000, 010000, 100000.

[0136] The embodiments of the present invention have flexible and diverse ways of performing feature transformation on randomness data. Based on the specific data of the extracted key features and the operation requirements of different scenarios, different feature transformation methods can be adopted.

[0137] Based on any of the above optional embodiments, after respectively performing feature transformation on each data and converting it into an encoded number composed of 0 and 1 through one-hot encoding, it specifically includes:

[0138] Based on all the data, obtain N numerical segments according to the first preset rule, sort the N numerical segments by numerical size, and obtain the sorting positions of each numerical segment, where N is an integer greater than 0;

[0139] Match each data with the N numerical segments according to the second preset rule, so that each data matches a numerical segment, and use the sorting position of the matched numerical segment as the feature transformation number of each data;

[0140] Perform one-hot encoding on the feature transformation number of each data to obtain an encoded number composed of 0 and 1.

[0141] Figure 2 This is a schematic diagram of numerical segmentation in the embodiments of the present invention. Assume that there are a total of 1150 data for which feature transformation is performed in the embodiments of the present invention, that is, 1150 data are obtained. Then, segmentation is performed according to a total of 1150 data, and the segmentation method is the first preset rule; the first preset rule can be to segment according to the numerical size of the 1150 data, or to segment according to the number of the 1150 data, or to segment by mixing numerical size and the number of data. The specific segmentation method can be determined according to requirements, and the embodiments of the present invention do not limit this.

[0142] As Figure 2 shown, it is assumed that in the embodiment of the present invention, after segmenting 1150 randomness data in the example according to the first preset rule, N numerical segments are obtained. After these N numerical segments are sorted by size, the sorting positions are 1, 2, …, N respectively.

[0143] After sorting the numerical segments, the embodiment of the present invention "assigns seats" to all the randomness data according to the second preset rule, where the second preset rule is a matching method corresponding to the first preset rule, that is: if the first preset rule is to segment according to the numerical size, the second preset rule is to match according to the numerical size; if the first preset rule is to segment according to the number of data, the second preset rule is to match according to the numbering of the number of data; if the first preset rule is to segment according to a mixture of numerical size and the number of data, the second preset rule is to match according to a mixture of numerical size and the number of data.

[0144] Please refer to Figure 2 , taking the first preset rule of segmenting according to the numerical size as an example, for the 1150 data in the above example, the 1150 data are respectively matched with N numerical segments. If the numerical values of 50 of the data are within the numerical range of segment 1, the sorting position 1 is used as the characteristic transformation number for these 50 data, that is, each of these 50 data obtains the characteristic transformation number 1; if the numerical values of 100 of the data are within the numerical range of segment 2, the sorting position 2 is used as the characteristic transformation number for these 100 data, that is, each of these 100 data obtains the characteristic transformation number 2; and so on. Each data will obtain a characteristic transformation number, which will not be elaborated here.

[0145] After segmenting and matching each data, each data obtains a characteristic transformation number. The embodiment of the present invention performs one-hot encoding on each characteristic transformation number to obtain an encoded number composed of 0 and 1; the one-hot encoding process in the following embodiments is the same and will not be elaborated hereafter.

[0146] It should be noted that since the word randomness in the embodiments of the present invention can include the word randomness of multiple languages, such as English word randomness, Chinese pinyin randomness, etc., the feature transformation and one-hot encoding can be performed on the English word randomness, Chinese pinyin randomness, and the first numerical data respectively. The English word randomness and Chinese pinyin randomness consist of different key feature strings, namely English word randomness and Chinese pinyin randomness, and the first numerical data also includes data with different key features. Therefore, in actual implementation, different key features can be processed separately, that is, the data of one feature transformation and one-hot encoding belongs to one key feature. For example, all of the above 1150 data are the quantity of metadata, or all of the 1150 data are the English word randomness calculated from the application package name, or all of the 1150 data are the Chinese pinyin randomness calculated from the application package name, and so on. Performing feature transformation and one-hot encoding according to the same key feature can generalize the data of this key feature, thereby improving the generalization ability of the AI model. The following feature transformations can all apply this strategy and will not be elaborated hereafter.

[0147] In the embodiments of the present invention, feature transformation is performed on the data, which can classify and discretize the data with similar characteristics (for example, the values are different but very close). This can improve the generalization ability of the AI model during training. Preferably, the embodiments of the present invention provide three feature transformation methods, namely, feature transformation by value, by the number of data, and by a combination of value and the number of data, as follows:

[0148] The first feature transformation method: The first preset rule is to segment by the numerical value, and the second preset rule is to match by the numerical value.

[0149] The second feature transformation method: The first preset rule is to segment by the number of data, and the second preset rule is to match by the serial number of the number of data.

[0150] The third feature transformation method: The first preset rule is to segment by a combination of the numerical value and the number of data, and the second preset rule is to match by a combination of the numerical value and the number of data.

[0151] The following will describe each feature transformation method in detail.

[0152] Based on any of the above optional embodiments, after each data is respectively subjected to feature transformation, through one-hot encoding, it is converted into an encoded number composed of 0 and 1, which specifically includes:

[0153] Taking the maximum and minimum values among all the data as the numerical range, dividing the numerical range into N parts to obtain N numerical segments, sorting the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; for each data, if the value of the data falls within the numerical range of a numerical segment S, then taking the sorting position of the numerical segment S as the feature transformation number of the data; performing one-hot encoding on the feature transformation numbers of each data to obtain the encoded numbers composed of 0 and 1.

[0154] This embodiment is the first feature transformation method, which is suitable for the scenario where the values of the data are continuously distributed. The following is an example to illustrate. Suppose the 1150 data in the above example have continuously distributed values, the maximum value is 1, and the minimum value is 0; take N = 10, the numerical range is 1000, divide the numerical range 1000 into 10 equal segments, and the numerical ranges of each segment are [0, 0.1], [0.1, 0.2], [0.2, 0.3], [0.3, 0.4], [0.4, 0.5], [0.5, 0.6], [0.6, 0.7], [0.7, 0.8], [0.8, 0.9], [0.9, 1] respectively. Sorting by the numerical range size of the numerical range, the corresponding sorting positions of the above numerical segments are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 respectively.

[0155] Match the 1150 data with the above 10 numerical segments respectively. If the value of a data a is 0.23 and falls within the segment [0.2, 0.3], then take the sorting position 3 of the segment [0.2, 0.3] as the feature transformation number of a; and so on. Match all the data by segments respectively, then each data will obtain a feature transformation number in the range of 1 to 10. Thus, the 1150 data are divided into 10 parts, one part obtains the feature transformation number 1, one part obtains the feature transformation number 2, and so on, and one part obtains the feature transformation number 10.

[0156] It should be noted that the sorting of the N numerical segments can be in ascending order or in descending order, and the sorting method during malicious application detection only needs to be consistent with the sorting method during AI model training.

[0157] In addition, for the value at the boundary point of the numerical segment, it can be matched to the smaller numerical segment or the larger numerical segment according to the agreed rules, and the embodiments of the present invention do not limit this.

[0158] In addition, in the above example, the numerical range 1000 is divided into 10 equal segments. If non-uniform segmentation is performed using specific rules, it also falls within the protection scope of the embodiments of the present invention.

[0159] Perform one-hot encoding on the data for which the feature transformation numbers have been obtained. For example, if the feature transformation number of data a is 1 and the total number of feature transformation numbers is 10, the one-hot encoding is 1000000000; if the feature transformation number of value-type data b is 2, the one-hot encoding is 0100000000; and so on.

[0160] Based on any of the above optional embodiments, after respectively performing feature transformation on each data, through one-hot encoding, it is converted into an encoded number composed of 0 and 1, specifically including:

[0161] Segment by the total number of all data to obtain N numerical segments, and sort the N numerical segments by numerical size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; sort each data by size to obtain the sorting number of each data; use the sorting position of the numerical segment corresponding to the sorting number of each data as the feature transformation number of each data; perform one-hot encoding on the feature transformation number of each data, so as to obtain an encoded number composed of 0 and 1.

[0162] This embodiment is the second feature transformation method, which is suitable for the scenario where the values of the data are discretely distributed. The following is an example to illustrate. Assume the 1150 data in the above example, whose values are discretely distributed. If the first feature transformation method is used, there may be no matching data in some numerical segments, and there may be a large number of matching data in some numerical segments, which will have an adverse impact on the data operation of the AI model.

[0163] To avoid the above situation, the second feature transformation method uses the number of data for segmentation and matching. Assume the 1150 data in the above example, take N = 10, then 10 numerical segments are obtained, and each numerical segment can match 115 data; sort the 1150 data by numerical size, then the 1st to 115th sorted data match numerical segment 1 and obtain the feature transformation number 1, the 116th to 230th data match numerical segment 2 and obtain the feature transformation number 2, and so on, the 1036th to 1150th data match numerical segment 10 and obtain the feature transformation number 10. The one-hot encoding for the data for which the feature transformation numbers have been obtained is the same as that in the above embodiment, and will not be elaborated here.

[0164] It should be noted that in the above example, the number of data 1150 is evenly divided into 10 segments. If non-uniform segmentation is performed using specific rules, it also falls within the protection scope of the embodiments of the present invention.

[0165] Based on any of the above optional embodiments, after respectively performing feature transformation on each data, through one-hot encoding, it is converted into an encoded number composed of 0 and 1, specifically including:

[0166] Part of all the data is segmented according to the number of data to obtain n numerical segments, where n is an integer greater than 0; part of the data segmented according to the number of data is eliminated, and the numerical interval is divided into Nn parts with the maximum and minimum values ​​of the remaining data as the numerical interval to obtain Nn numerical segments, where N is an integer greater than n; the N numerical segments are sorted according to the numerical size to obtain the ranking rank of each numerical segment; each data is sorted according to the size, and matched with the N numerical segments respectively, and the ranking rank of the matched numerical segments is used as the feature transformation number of the data; the feature transformation number of each data is uniquely encoded to obtain a coded number composed of 0 and 1.

[0167] This embodiment is the third feature transformation method, which is suitable for scenarios where the data values ​​have both discrete distribution and continuous distribution, and is explained below by example. It should be noted that the order of the text description of the third feature transformation method is not used to limit the order of the actual implementation steps; in fact, all methods of segmenting by mixing the value size and the number of data, and matching by mixing the value size and the number of data, all fall within the protection scope of the embodiments of the present invention.

[0168] Assuming the 1150 data in the above example, the 1150 data are sorted by numerical value, wherein the first and last data segments are relatively discrete, and the data in the middle are relatively continuous; then the first data segment, i.e., x% of the total data, is taken as the first numerical segment, and the sorting position is 1, and the last data segment, i.e., y% of the total data, is taken as the Nth numerical segment, and the sorting position is N. Take N=10, x=10, y=20, 1150*10%=115, 1150*20%=230, then the first numerical segment can match 115 data, and the Nth numerical segment can match 230 data; after sorting the 1150 data by numerical value, the 1st to 115th data are matched to the first numerical segment, and the feature transformation number 1 is obtained, and the 921st to 1150th data are matched to the Nth numerical segment, and the feature transformation number 10 is obtained.

[0169] The maximum and minimum values ​​of the 116th to 920th data among the 1150 data sorting positions are taken as the numerical interval, and the numerical interval is divided into 10-2 segments, i.e., 8 numerical segments, by an average or non-average method. The sorting positions are 2, 3, 4, 5, 6, 7, 8, and 9 respectively; each data from the 116th to 920th data is matched with the numerical segments 2 to 9 respectively. If the value of the data is within the numerical range of a numerical segment S, the sorting position of the numerical segment S is taken as its characteristic transformation number.

[0170] The one-hot encoding of the data whose feature transformation numbers have been obtained is the same as that in the above embodiment and will not be repeated here.

[0171] It should be noted that in the above example, the first and last segments of the sorted data are segmented according to the number of data, and the middle part is segmented according to the numerical value range. It is also possible to segment the first and last segments according to the numerical value range according to the needs of time applications, and segment the middle part according to the number of data, and so on. As long as the data segmentation includes segmentation according to the numerical value range and segmentation according to the number of data, all fall within the protection scope of the embodiments of the present invention.

[0172] The embodiments of the present invention perform feature transformation on the extracted key features and then perform one-hot encoding. This is based on the data itself. During the training of the AI model, while ensuring that the information contained in the data can be maximally extracted, the final effect of the AI model is guaranteed; during the detection of malicious application programs, the accuracy of the detection is maximally guaranteed.

[0173] Based on any of the above optional embodiments, the AI model network structure of the embodiments of the present invention includes, but is not limited to, convolutional neural networks, recurrent neural networks, deep learning networks, machine learning networks, etc. It mainly includes four parts: an input layer, a factorization machine layer, a hidden layer, and an output layer. Figure 3 is a schematic diagram of the network structure of the AI model according to the embodiments of the present invention. As Figure 3 shown, the network structure of the AI model includes four parts: an input layer (Input Layers), a factorization machine layer (Factorization Machines), a hidden layer (Hidden Layers), and an output layer (Output Layers);

[0174] The input layer is used to receive digital feature vectors and application program malicious labels as input data; the factorization machine layer is used to extract low-order features from the input data and perform calculations based on the low-order features; the hidden layer is used to extract high-order features from the input data, calculate the malicious features of the application program according to the high-order features, and segment the malicious features of application programs with different maliciousness from a high-dimensional space; the output layer is used to merge the calculation results of the factorization machine layer and the calculation results of the hidden layer, and output the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application.

[0175] Figure 4 is a schematic diagram of the training process of the AI model according to the embodiments of the present invention. As Figure 4 shown, the AI model according to the embodiments of the present invention is trained by the following method:

[0176] 400, obtain application program installation package samples and the malicious labels of the application program installation package samples;

[0177] 401. Analyze the second target file in the application installation package sample, and extract the second key features in the second target file; the second key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension;

[0178] 402. Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a digital feature vector;

[0179] 403. Input the second digital feature vector and the malicious label of the application installation package sample into the established AI model, and train the AI model to obtain an AI model that meets the expected requirements.

[0180] It should be noted that the second key features and the second digital feature vector in the embodiments of the present invention are only for making a noun distinction from the key features and the digital feature vector mentioned above, and the "second" has no actual meaning; the actual meanings of the second key features and the key features are exactly the same, and the actual meanings of the second digital feature vector and the digital feature vector are exactly the same.

[0181] In the embodiments of the present invention, the application installation package sample data and the malicious label data obtained in step 400 can be the results of manual analysis, and the samples are collected and the maliciousness is judged manually. The embodiments of the present invention collect a large number of complete Android platform application programs, analyze the characteristics of application malicious trojans, etc. manually, and then extract the rules for judging the maliciousness of the application programs. After that, based on the rules, the maliciousness of the collected Android platform application programs is judged to obtain the malicious label corresponding to the application installation package sample.

[0182] As described above, the processing procedures in the AI model training phase and the malicious application detection phase of the embodiments of the present invention are exactly the same, including calculating the word randomness for some key features and obtaining numerical data (further performing feature transformation) for the remaining key features; if during AI model training, the randomness data and numerical data of some key features are directly concatenated into a digital feature vector as the input data during AI model training, then in the malicious application detection phase, the randomness data and numerical data of some key features are also directly concatenated into a digital feature vector as the input data during AI model training; if during AI model training, the randomness data of some key features and some numerical data (the first numerical data) are subjected to feature transformation and one-hot encoding, and then concatenated with the remaining numerical data (the second numerical data) into a digital feature vector as the input data during AI model training, then in the malicious application detection phase, the randomness data of some key features and some numerical data (the first numerical data) are also subjected to feature transformation and one-hot encoding, and then concatenated with the remaining numerical data (the second numerical data) into a digital feature vector as the input data during AI model training; the processing steps of all optional embodiments in the AI model training phase and the malicious application detection phase are exactly the same, so as to ensure the detection effect of the AI model.

[0183] The processing method of step 401 of the AI model training method in the embodiments of the present invention is exactly the same as that of step 101 of the malicious application detection method based on features and all optional embodiments of step 101; if there are multiple different optional embodiments, keeping the optional embodiments adopted by step 101 and step 401 consistent can achieve the effect of malicious application detection, which will not be elaborated here.

[0184] The processing method of step 402 of the AI model training method in the embodiments of the present invention is exactly the same as that of step 102 of the malicious application detection method based on features and all optional embodiments of step 102; if there are multiple different optional embodiments, keeping the optional embodiments adopted by step 102 and step 402 consistent can achieve the effect of malicious application detection, which will not be elaborated here.

[0185] The difference between step 403 of the AI model training method in the embodiments of the present invention and step 103 of the malicious application detection method based on features is that the input data of step 403 includes, in addition to the digital feature vector concatenated by the randomness of the key features, the malicious label of the application installation package sample, and the AI model is trained by the digital feature vector and the malicious label of the application installation package sample, so that the AI model has the ability to identify malicious applications based on the digital feature vector.

[0186] The factorization machines layer of the AI model in the embodiments of the present invention is used to extract low-order features in the data. The maliciousness of an APP is often a series of actions. Only by cooperating with each other in terms of behavior, permissions, content, etc. can malicious behavior be completed. Therefore, the combination of features can often reflect the maliciousness of the APP to a great extent. The factorization machine mainly extracts feature combinations through the inner product of latent variables for each dimension of features to ensure that the model can preserve the combination information between features.

[0187] Preferably, the hidden layers of the AI model in the embodiments of the present invention is a feed-forward neural network, including hidden units and sigmoid functions; there are 6 hidden layers in the hidden layer, and the number of neuron nodes in each layer ranges from 32 to 256. Specifically, the number of neuron nodes in each layer can be 64, 128, 256, 128, 64, or 32, which is used to extract high-order features in the data and segment APP features with different maliciousness from a high-dimensional space.

[0188] The output layer of the AI model in the embodiments of the present invention combines the forward calculation results of the factorization machines layer and the hidden layer, and the final number of output nodes is 2 or 1: if it is 2, they are respectively the probability that this sample is a non-malicious application and the probability that it is a malicious application; if it is 1, it is the probability that this sample is a non-malicious application or the probability that it is a malicious application.

[0189] Preferably, by setting the probability thresholds for malicious applications and non-malicious applications, the detection requirements for low false alarm scenarios can be met.

[0190] Finally, the model is trained by backpropagation; taking cross-validation as a means and accuracy, recall rate, precision, and F-value as criteria, considering the influence of time factors and business background, the parameters, number of iterations, model structure, etc. of the AI model are adjusted, and finally an AI model that meets the expected requirements is obtained.

[0191] Based on any of the above optional embodiments, step 403 of inputting the second digital feature vector and the malicious label of the application program installation package sample into the built AI model and training the AI model specifically includes:

[0192] 403.1, converting the malicious label of the application program installation package sample into the malicious label of the second digital feature vector;

[0193] 403.2, input the second digital feature vector and the malicious label of the second digital feature vector into the established AI model to train the AI model.

[0194] In the embodiment of the present invention, the malicious label of the application installation package sample is the result of tagging the maliciousness of the APP after analyzing a large number of mainstream malicious software on the Android platform. An application installation package sample corresponds to a digital feature vector and also corresponds to a malicious label. Therefore, the digital feature vector of the application installation package sample can be corresponded to the malicious label, and the malicious label of the application installation package sample can be converted into the malicious label of the digital feature vector, that is, one digital feature vector corresponds to one malicious label.

[0195] In this way, one input data of the AI model in the embodiment of the present invention is the digital feature vector, and the other input data is the malicious label corresponding to the digital feature vector, and the AI model is trained.

[0196] In the embodiment of the present invention, the digital feature vector is corresponded to the malicious label. By tagging the digital feature vector, the generalization ability of the AI model can be improved.

[0197] Based on any of the above optional embodiments, the conversion of the malicious label of the application installation package sample in step 403.1 into the malicious label of the second digital feature vector specifically includes:

[0198] Based on the malicious labels of all application installation package samples, use the malicious label of any one application installation package sample as the malicious label of the second digital feature vector corresponding to the any one application installation package sample;

[0199] Based on all the second digital feature vectors and their corresponding malicious labels, remove the duplicates of the data with exactly the same second digital feature vectors and their corresponding malicious labels, and obtain the de-duplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

[0200] In the embodiments of the present invention, the malicious labels of the application installation package samples are converted into malicious labels of digital feature vectors. Inevitably, there will be data where the digital feature vectors are exactly the same as the malicious labels. For example, application A is malicious after analysis, application B is malicious after analysis, and application C is malicious after analysis; assume that after applications A, B, and C extract key features and perform feature transformation, the resulting digital feature vectors are exactly the same, assumed to be 0010000000, and the malicious labels of A, B, and C are all malicious, assumed to be 1. Then there will be three identical data entries (0010000000, 1). The embodiments of the present invention remove duplicates from the data where the digital feature vectors and their corresponding malicious labels are exactly the same, which can significantly reduce the data volume and relieve the data processing pressure on the AI model. After testing in a laboratory scenario, after data deduplication, the data volume is reduced by more than one order of magnitude approximately.

[0201] Furthermore, for multiple identical digital feature vectors, if the corresponding malicious labels include malicious and non-malicious, and if the number of malicious ones is greater than the number of non-malicious ones, then this digital feature vector is labeled as malicious; otherwise, it is labeled as non-malicious.

[0202] For example, for applications A, B, C, D, and E, after extracting key features and performing feature transformation, the resulting digital feature vectors are exactly the same, assumed to be 0010000000. The malicious labels of A, B, and C are all malicious, assumed to be 1, and the malicious labels of D and E are both non-malicious, assumed to be 0. Then the number of digital feature vectors 0010000000 corresponding to malicious label 1 is greater than the number corresponding to malicious label 0, so the digital feature vector 0010000000 is labeled as 1.

[0203] In this embodiment, according to the number of malicious and non-malicious corresponding malicious labels of the digital feature vectors, the digital feature vectors are de-duplicated and re-labeled, further reducing the data volume, relieving the data processing pressure on the AI model, and at the same time further improving the generalization ability of the AI model.

[0204] In summary, the embodiments of the present invention first extract several key features such as application package name, program name, developer information, language information, behaviors, and permissions from three dimensions of behaviors, permissions, and content based on the analysis of Android platform applications; secondly, calculate the randomness of some key features, obtain numerical values for some key features, and further perform data preprocessing such as feature transformation on the randomness data and some numerical data; thirdly, use a model redesigned based on the DeepFM algorithm in the deep learning method to perform multiple rounds of model training, and make the model effect reach the expected goal through the evaluation strategy of the model effect; finally, extract features of applications with unknown maliciousness and predict through the model, and then complete the determination of the maliciousness of the applications.

[0205] Among them, the key features are the cornerstone of AI model training and the basis for malicious application detection. The key features include information such as application-specific behaviors, permissions, sensitive strings, certificates, etc. Malicious applications have strong differentiations from normal applications in these specific dimensions, which are important bases for judging the maliciousness of applications. By verifying and parsing the application installation package, a large number of keywords are obtained. Applying the randomness algorithm to a large number of keywords, the randomness data of the keywords is obtained. All the randomness data is concatenated to form a vector, which is an important basis for judging whether the application is suspected of being batch-generated by malicious tools and whether these batch-generated applications carry viruses.

[0206] Furthermore, in the embodiments of the present invention, by performing feature transformation on the randomness data and some numerical data, from the perspective of the model, while ensuring that the information contained in the data can be maximally extracted, the final effect of the AI model is ensured.

[0207] Model design is the key point for the success of the AI model. Designing an appropriate model can maximize the conversion of information in the data into knowledge, and then solidify the knowledge for use in judging the maliciousness of applications.

[0208] In summary, the embodiments of the present invention perform static feature extraction and data statistical features on a large number of Android platform applications, redesign the algorithm based on the DeepFM algorithm in the deep learning method, and then perform model training on the data after feature extraction. Finally, an AI model for detecting malicious applications on the Android platform is obtained. By using the AI model to detect malicious applications, it does not rely on the update of the malicious feature library, has strong learning ability, and provides a new direction for the malicious detection of Android platform applications. The present invention solves the problems existing in the traditional malicious application detection based on manual extraction of rules, such as difficult rule extraction, low coverage, poor scalability, and easy to be bypassed, and the embodiments of the present invention have higher accuracy and timeliness.

[0209] Figure 5 The malicious application detection device based on features according to the embodiments of the present invention includes a key feature extraction module 501, a randomness acquisition module 502, and a malicious application detection module 503:

[0210] The key feature extraction module 501 parses the target file in the application installation package to be detected, and extracts the key features in the target file. The key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension;

[0211] The randomness acquisition module 502 acquires the word randomness of some key features, acquires the numerical data of the remaining key features, and splices the word randomness and the numerical data into a digital feature vector;

[0212] The malicious application detection module 503 inputs the digital feature vector into a trained AI model to obtain a maliciousness detection result for the application to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0213] The feature-based malicious application detection device according to an embodiment of the present invention is used to execute Figure 1 the technical solution of the feature-based malicious application detection method embodiment shown, and its implementation principle and technical effect are similar, which will not be elaborated here.

[0214] Figure 6 It is a schematic diagram of the electronic device framework according to an embodiment of the present invention. Please refer to Figure 6 , an embodiment of the present invention provides an electronic device, including: a processor 610, a communication interface 620, a memory 630, and a bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following methods, including: parsing the target file in the application installation package to be detected, extracting the key features in the target file, and the key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension; acquiring the word randomness of some key features, acquiring the numerical data of the remaining key features, and splicing the word randomness and the numerical data into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result for the application to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0215] An embodiment of the present invention discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments. For example, it includes: parsing a target file in an application program installation package to be detected, extracting key features in the target file, where the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; obtaining the word randomness of some key features, obtaining the numerical data of the remaining key features, and splicing the word randomness and the numerical data into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0216] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions. The computer instructions cause the computer to execute the methods provided in the above-mentioned method embodiments. For example, it includes: parsing a target file in an application program installation package to be detected, extracting key features in the target file, where the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; obtaining the word randomness of some key features, obtaining the numerical data of the remaining key features, and splicing the word randomness and the numerical data into a digital feature vector; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application program to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application program corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0217] Those of ordinary skill in the art can understand that implementing the above device embodiment or method embodiment is merely illustrative. The processor and the memory may or may not be physically separated components, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0218] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a USB flash drive, a mobile hard disk, ROM / RAM, a magnetic disk, an optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0219] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A feature-based malicious application detection method, characterized in that, it includes: Analyze the target file in the installation package of the application to be detected, and extract the key features in the target file. The key features include at least one dimension information of the behavior dimension, the permission dimension, and the content dimension; Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a digital feature vector; Input the digital feature vector into the trained AI model to obtain the maliciousness detection result of the application to be detected; the AI model is trained according to the input digital feature vector and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application; The obtaining of the word randomness of the partial key features specifically includes: According to the alphabetical order of the partial key feature strings, sequentially obtain the letter adjacent frequencies of any two adjacent letters of the partial key feature strings; Based on the letter adjacent frequencies of any two adjacent letters, obtain the word randomness of the key features; Wherein, the letter adjacent frequency is to regard the string as a word of a preset language text, and obtain the adjacent frequency between two letters through the word randomness calculation rule of the preset language text; The word randomness is obtained through the letter adjacent frequency; The obtaining of the numerical data of the remaining key features specifically includes: Based on the key features in the content dimension of the remaining key features, convert the extracted key features into first numerical data, and the first numerical data is the numerical value represented by the key features; Based on the key features in the behavior dimension and the permission dimension of the remaining key features, convert the extracted key features into second numerical data, and the second numerical data is the number 0 or 1; The splicing of the word randomness and the numerical data into a digital feature vector specifically includes: Based on each data in the word randomness and the first numerical data, respectively perform feature transformation on each data, and then through one-hot encoding, convert it into an encoded number composed of 0 and 1; Splice all the encoded numbers after the randomness data and the first numerical data are converted, and the second numerical data, into a digital feature vector in a preset order.

2. The method according to claim 1, characterized in that, the behavior dimension includes the behavior information when the application is running; the permission dimension includes the permission information required by the application when performing specific behaviors; the content dimension includes at least one of the following information: application package name, program name, developer information, language information, total file size of the application, number of files included in the application, size of a specific file in the application, number of specific components in the application, number of metadata, number of resource strings, number of supported languages, and application reinforcement information.

3. The method according to claim 2, characterized in that, the step of respectively performing feature transformation on each data and then through one-hot encoding to convert it into an encoded number composed of 0 and 1 specifically includes: Based on all the data, obtain N numerical segments according to the first preset rule, and sort the N numerical segments according to the numerical value size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; Match each data with the N numerical segments according to the second preset rule, so that each data matches a numerical segment, and use the sorting position of the matched numerical segment as the feature transformation number of each data; Perform one-hot encoding on the feature transformation number of each data to obtain a coded number composed of 0s and 1s.

4. The method according to claim 2 or 3, wherein, After respectively performing feature transformation on each data, through one-hot encoding, converting it into a coded number composed of 0s and 1s, specifically includes: Taking the maximum value and the minimum value in all the data as the numerical interval, dividing the numerical interval into N parts to obtain N numerical segments, and sorting the N numerical segments according to the numerical value size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; For each data, if the value of the data is within the numerical range of a numerical segment S, then use the sorting position of the numerical segment S as the feature transformation number of the data; Perform one-hot encoding on the feature transformation number of each data to obtain a coded number composed of 0s and 1s.

5. The method according to claim 2 or 3, wherein, After respectively performing feature transformation on each data, through one-hot encoding, converting it into a coded number composed of 0s and 1s, specifically includes: Segment according to the total number of all the data to obtain N numerical segments, and sort the N numerical segments according to the numerical value size to obtain the sorting positions of each numerical segment, where N is an integer greater than 0; Sort each data according to the size to obtain the sorting number of each data; Use the sorting position of the numerical segment corresponding to the sorting number of each data as the feature transformation number of each data; Perform one-hot encoding on the feature transformation number of each data to obtain a coded number composed of 0s and 1s.

6. The method according to claim 2 or 3, wherein, After respectively performing feature transformation on each data, through one-hot encoding, converting it into a coded number composed of 0s and 1s, specifically includes: Segment a part of all the data according to the number of data to obtain n numerical segments, where n is an integer greater than 0; Eliminate the part of the data segmented according to the number of data, and take the maximum value and the minimum value in the remaining data as the numerical interval, divide the numerical interval into N - n parts to obtain N - n numerical segments, where N is an integer greater than n; Sort the N numerical segments according to the numerical value size to obtain the sorting positions of each numerical segment; Sort each data according to the size, and respectively match it with the N numerical segments, and use the sorting position of the matched numerical segment as the feature transformation number of the data; Perform one-hot encoding on the feature transformation number of each data to obtain a coded number composed of 0s and 1s.

7. The method according to any one of claims 1 - 3, wherein, The network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method: Obtain application installation package samples and the malicious labels of the application installation package samples; Parse the second target file in the application installation package sample, and extract the second key features in the second target file, where the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a second digital feature vector; Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.

8. The method according to claim 4, wherein, The network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method: Obtain application installation package samples and the malicious labels of the application installation package samples; Parse the second target file in the application installation package sample, and extract the second key features in the second target file, where the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a second digital feature vector; Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.

9. The method according to claim 5, wherein, The network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method: Obtain application installation package samples and the malicious labels of the application installation package samples; Parse the second target file in the application installation package sample, and extract the second key features in the second target file, where the key features include at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a second digital feature vector; Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.

10. The method according to claim 6, wherein, The network structure of the AI model includes four parts: an input layer, a decomposition machine layer, a hidden layer, and an output layer; the AI model is trained by the following method: Obtain application installation package samples and the malicious labels of the application installation package samples; Analyze the second target file in the application installation package sample, and extract the second key features in the second target file. The key features include at least one dimension information among the behavior dimension, permission dimension, and content dimension; Obtain the word randomness of some key features, obtain the numerical data of the remaining key features, and splice the word randomness and the numerical data into a second digital feature vector; Input the second digital feature vector and the malicious label of the application installation package sample into the built AI model, and train the AI model to obtain an AI model that meets the expected requirements.

11. The method according to claim 7, wherein, The step of inputting the second digital feature vector and the malicious label of the application installation package sample into the built AI model and training the AI model specifically includes: Convert the malicious label of the application installation package sample into the malicious label of the second digital feature vector; Input the second digital feature vector and the malicious label of the second digital feature vector into the built AI model, and train the AI model.

12. The method according to any one of claims 8-10, wherein, The step of inputting the second digital feature vector and the malicious label of the application installation package sample into the built AI model and training the AI model specifically includes: Convert the malicious label of the application installation package sample into the malicious label of the second digital feature vector; Input the second digital feature vector and the malicious label of the second digital feature vector into the built AI model, and train the AI model.

13. The method according to claim 11, wherein, The step of converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes: Based on the malicious labels of all application installation package samples, use the malicious label of any one application installation package sample as the malicious label of the second digital feature vector corresponding to the any one application installation package sample; Based on all the second digital feature vectors and their corresponding malicious labels, remove duplicates from the data where the second digital feature vectors and their corresponding malicious labels are exactly the same, and obtain the de-duplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

14. The method according to claim 12, wherein, The step of converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes: Based on the malicious labels of all application installation package samples, use the malicious label of any one application installation package sample as the malicious label of the second digital feature vector corresponding to the any one application installation package sample; Based on all the second digital feature vectors and their corresponding malicious labels, remove duplicates from the data where the second digital feature vectors and their corresponding malicious labels are exactly the same, and obtain the de-duplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

15. An electronic device, characterized in that, comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 14 by invoking the program instructions.

16. A non-transitory computer-readable storage medium, characterized in that, the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Word vector- and depth neural network-based Android malicious code detection method

    CN108959924A

  • Malicious program detection method and device and related application

    CN109726554A

  • Malicious application program detection method and equipment based on randomness degree

    CN113971281A

  • Malicious application program detection method and equipment based on AI model

    CN113971282A