A malicious application detection method and device based on AI model

Through the malicious application detection method based on AI model, static information of the application is extracted and processed, and the problem of difficult detection of new malicious applications in the existing technology is solved, and higher detection accuracy and timeliness are achieved.

CN113971282BActive Publication Date: 2025-06-06WUHAN ANTIY MOBILE SECURITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010721509.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-24
Publication Date
2025-06-06
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

Existing mobile security software relies on the update of malicious feature library, making it difficult to effectively detect new malicious applications, and is difficult to extract rules, has low coverage, poor scalability, and is easily bypassed.

Method used

Using a malicious application detection method based on AI model, we parse the target files in the application installation package, extract static information, and process the information into a digital feature vector through feature transformation, and input the trained AI model to obtain malicious detection results.

Benefits of technology

Improve the accuracy and timeliness of malicious program detection, avoid the limitation of relying on malicious feature library updates, and enhance the detection capabilities of new malicious applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971282B_ABST
    Figure CN113971282B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and device for detecting malicious applications based on an AI model. The method includes parsing a target file in an installation package of an application to be detected, extracting static information from the target file, wherein the static information includes at least one dimension of information in a behavior dimension, a permission dimension, and a content dimension; processing the static information into a digital feature vector by feature transformation, wherein the digital feature vector is composed of 0 and 1; inputting the digital feature vector into a trained AI model to obtain a malicious detection result for the application to be detected; the AI ​​model is trained based on the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that the application is a non-malicious application. It solves the problems of difficult rule extraction, low coverage, poor scalability, and easy bypass in traditional malicious application detection, and has higher accuracy and timeliness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of mobile network security technology, and in particular to a malicious application detection method and device based on an AI model. Background Art

[0002] At present, many security vendors have also invested in the field of mobile security. The basic principle of these software for antivirus is to confirm the intrusion behavior by matching the known malicious Trojan features, and to actively defend with firewalls, dynamic monitoring and other methods. However, the disadvantage is that they rely on the update of the malicious feature library and have a weak ability to learn new malicious detection.

[0003] However, new malicious applications emerge in an endless stream, and their maliciousness varies. Relying on malicious feature libraries for malicious detection will make it difficult to achieve ideal security protection effects due to untimely updates of malicious feature libraries. Summary of the invention

[0004] In view of the problems existing in the prior art, the embodiments of the present invention provide a malicious application detection method and device based on an AI model. Based on the analysis of a large number of mainstream malware on the Android platform, the attack intentions and means of malware on the Android platform are summarized, and AI detection of application malice is implemented through a deep learning algorithm, providing a new direction for malicious detection of Android platform applications.

[0005] In a first aspect, an embodiment of the present invention provides a malicious application detection method based on an AI model, comprising:

[0006] Parsing a target file in an installation package of the application to be detected, and extracting static information in the target file, wherein the static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;

[0007] Processing the static information into a digital feature vector by feature transformation, wherein the digital feature vector is composed of 0 and 1;

[0008] The digital feature vector is input into a trained AI model to obtain a maliciousness detection result for the application to be detected; the AI ​​model is trained based on the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that the application is a non-malicious application.

[0009] Specifically, the behavior dimension includes behavior information when the application is running;

[0010] The permission dimension includes permission information required by the application when performing a specific behavior;

[0011] The content dimension includes at least one of the following information: the total file size of the application, the number of files included in the application, the size of specific files in the application, the number of specific components in the application, and application hardening information.

[0012] Further, the processing of the static information into a digital feature vector by feature transformation specifically includes:

[0013] Based on each dimension data in the static information, the numerical data in each dimension data is subjected to feature transformation, and then converted into a coded number consisting of 0 and 1 through one-hot encoding;

[0014] Convert the Boolean data in each dimension data into a number 0 or 1 respectively;

[0015] The numbers after all numerical data and Boolean data conversion are concatenated into a digital feature vector composed of 0 and 1 in a preset order.

[0016] Furthermore, the feature transformation of the numerical data in each dimension data is performed and then converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including:

[0017] Based on the numerical data in each dimension data, obtain N numerical segments according to a first preset rule, and sort the N numerical segments according to numerical values ​​to obtain a ranking of each numerical segment, where N is an integer greater than 0;

[0018] Matching each numerical data in each dimension data with the N numerical segments according to a second preset rule, so that each numerical data matches one numerical segment, and using the ranking of the matched numerical segments as the characteristic transformation number of each numerical data;

[0019] The feature transformation digits of each numerical data are uniquely encoded to obtain encoded digits consisting of 0 and 1.

[0020] Furthermore, the feature transformation of the numerical data in each dimension data is performed and then converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including:

[0021] Based on the numerical data in each dimension data, the maximum value and the minimum value in the numerical data are used as the numerical interval, the numerical interval is divided into N parts to obtain N numerical segments, and the N numerical segments are sorted according to the numerical size to obtain the ranking of each numerical segment, where N is an integer greater than 0;

[0022] For each numerical data in each dimension data, if the value of the numerical data is within the numerical range of a numerical segment S, the ranking of the numerical segment S is used as the characteristic transformation number of the numerical data;

[0023] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0024] Furthermore, the feature transformation of the numerical data in each dimension data is performed and then converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including:

[0025] Based on the numerical data in each dimension data, segment the data according to the total number of the numerical data to obtain N numerical segments, and sort the N numerical segments according to the numerical values ​​to obtain the ranking of each numerical segment, where N is an integer greater than 0;

[0026] Sort the numerical data in each dimension data by size, and obtain the sorting number of each numerical data;

[0027] For each numerical data, the ranking of the numerical segment corresponding to the ranking number of each numerical data is used as the characteristic transformation number of each numerical data;

[0028] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0029] Furthermore, the feature transformation of the numerical data in each dimension data is performed and then converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including:

[0030] Based on the numerical data in each dimension data, segment part of the numerical data according to the number of data to obtain n numerical segments, where n is an integer greater than 0;

[0031] Eliminate some of the numerical data that are segmented according to the number of data, and use the maximum and minimum values ​​of the remaining data as the numerical interval, divide the numerical interval into Nn parts, and obtain Nn numerical segments, where N is an integer greater than n;

[0032] Sort the N numerical segments according to their numerical values ​​to obtain the ranking of each numerical segment;

[0033] Sort the numerical data in each dimension data by size, match them with N numerical segments respectively, and use the ranking of the matched numerical segments as the feature transformation number of the numerical data;

[0034] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0035] Furthermore, the network structure of the AI ​​model includes four parts: an input layer, a decomposition layer, a hidden layer, and an output layer; the AI ​​model is trained by the following method:

[0036] Obtaining application installation package samples and malicious labels of the application installation package samples;

[0037] Parsing a second target file in the application installation package sample, and extracting second static information in the second target file; the second static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;

[0038] Processing the second static information into a second digital feature vector by feature transformation, wherein the second digital feature vector consists of 0 and 1;

[0039] The second digital feature vector and the malicious label of the application installation package sample are input into the constructed AI model, and the AI ​​model is trained to obtain an AI model that meets the expected requirements.

[0040] Furthermore, the step of inputting the second digital feature vector and the malicious label of the application installation package sample into the constructed AI model to train the AI ​​model specifically includes:

[0041] Converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector;

[0042] The second digital feature vector and the malicious label of the second digital feature vector are input into the constructed AI model to train the AI ​​model.

[0043] Further, converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes:

[0044] Based on the malicious labels of all application installation package samples, taking the malicious label of any application installation package sample as the malicious label of the second digital feature vector corresponding to the any application installation package sample;

[0045] Based on all the second digital feature vectors and their corresponding malicious labels, data in which the second digital feature vectors and their corresponding malicious labels are completely identical are deduplicated to obtain the deduplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

[0046] In a second aspect, an embodiment of the present invention provides an electronic device, including:

[0047] at least one processor; and

[0048] at least one memory in communication with the processor, wherein:

[0049] The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the malicious application detection method based on the AI ​​model described in the first aspect of the embodiment of the present invention and the method described in any optional embodiment thereof.

[0050] In a third aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions execute the AI ​​model-based malicious application detection method described in the first aspect of the embodiment of the present invention and the method of any optional embodiment thereof.

[0051] The AI ​​model-based malicious application detection method provided by the embodiment of the present invention extracts static information of the target file in the installation package of the application to be detected, and processes the static information into a digital feature vector by feature transformation; based on the trained AI model, the digital feature vector is operated by the AI ​​model to obtain the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application. The embodiment of the present invention solves the problems of difficult rule extraction, low coverage, poor scalability, and easy bypass in the traditional malicious application detection based on manually extracted rules, and has higher accuracy and timeliness in malicious program detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0053] Figure 1 This is a flow chart of a malicious application detection method based on an AI model according to an embodiment of the present invention;

[0054] Figure 2 It is a numerical segmentation schematic diagram of an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of the network structure of the AI ​​model according to an embodiment of the present invention;

[0056] Figure 4A schematic diagram of the AI ​​model training process according to an embodiment of the present invention;

[0057] Figure 5 A malicious application detection device based on an AI model according to an embodiment of the present invention;

[0058] Figure 6 It is a schematic diagram of the framework of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] In view of the problems in the prior art, an embodiment of the present invention redesigns an algorithm based on a deep learning algorithm to obtain an artificial intelligence (AI) model; based on static information extraction and data statistical features of a large number of Android platform applications, the AI ​​model is trained to finally obtain an AI model for detecting malicious applications on the Android platform, wherein the training process includes: extracting static information from sample data (application installation package samples); performing feature transformation on the extracted static information to obtain a digital feature vector corresponding to the sample data; using the digital feature vector corresponding to the sample data and the malicious label corresponding to the sample data as input data of the AI ​​model to train the AI ​​model, and evaluating the effect through a model evaluation strategy to obtain an AI model that meets the expected requirements.

[0061] When performing malicious detection on an application, static information is extracted from the target file in the installation package of the application to be detected, and feature transformation is performed on the extracted static information to obtain a digital feature vector corresponding to the installation package of the application to be detected; the digital feature vector corresponding to the installation package of the application to be detected is used as input data of the AI ​​model, and after calculation, the AI ​​model outputs the probability that the installation package of the application to be detected is a malicious application and / or the probability that it is a non-malicious application.

[0062] In the embodiment of the present invention, the AI ​​model training stage and the malicious application detection stage both include static information extraction and feature transformation of the extracted static information, and their processing methods are exactly the same; in all optional schemes of the embodiment of the present invention, the scheme adopted in the malicious detection stage of the application is consistent with the scheme adopted in the AI ​​model training stage to obtain the expected detection effect.

[0063] The following is a detailed description of the malicious application detection method based on the AI ​​model described in the embodiment of the present invention from the perspective of the malicious application detection stage.

[0064] Figure 1 FIG. 1 is a flow chart of a malicious application detection method based on an AI model according to an embodiment of the present invention. Figure 1 The AI ​​model-based malicious application detection method shown includes:

[0065] 101, parsing a target file in an installation package of the application to be detected, and extracting static information in the target file, wherein the static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;

[0066] 102, processing the static information into a digital feature vector by feature transformation, wherein the digital feature vector consists of 0 and 1;

[0067] 103. Input the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application to be detected; the AI ​​model is trained based on the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that the application is a non-malicious application.

[0068] When the embodiment of the present invention performs malicious detection on an application, the installation package of the application to be detected is first obtained, the target file in the installation package of the application to be detected is parsed, and the static information in the target file is extracted. The static information includes at least one dimension information of the behavior dimension, the permission dimension, and the content dimension. Each dimension information may include one or more different information. For example, the behavior dimension may include behavior 1, behavior 2, behavior 3, etc., the permission dimension may include permission 1, permission 2, permission 3, permission 4, etc., and the content dimension may include content 1, content 2, content 3, etc.

[0069] Step 102 performs feature transformation on the extracted static information. The embodiment of the present invention converts different types of static information data into a unified digital feature vector, with the purpose of retaining more feature information while ensuring the generalization ability of the AI ​​model. The embodiment of the present invention mainly includes two data types, one is a numerical type formed by counting keywords or the number of components, and the other is a Boolean type that indicates whether a behavior or permission exists. Generally, the static information of the behavior dimension and the permission dimension is Boolean type data, and the static information of the content dimension is numerical type data, but unexpected situations are not excluded.

[0070] Step 103 inputs the digital feature vector composed of 0 and 1 into the pre-trained AI model, and obtains the maliciousness detection result of the application to be detected by calculation; the maliciousness detection result in the embodiment of the present invention refers to the probability that the application to be detected is a malicious application and / or the probability of a non-malicious application. The AI ​​model in the embodiment of the present invention can output the probability that the application is a malicious application, or output the probability that the application is a non-malicious application, or output the probability that the application is a malicious application and the probability of a non-malicious application at the same time.

[0071] Assume that the AI ​​model outputs the probability of malicious applications and the probability of non-malicious applications at the same time. If the former is greater than the latter, the application is non-malicious, otherwise the application is malicious. For example, the output is "[0.98654, 0.01346]", where 0.98654 is the probability of a non-malicious application and 0.01346 is the probability of a malicious application, indicating that this program is non-malicious, thereby completing the judgment of the maliciousness of the application.

[0072] The AI ​​model-based malicious application detection method described in the embodiment of the present invention extracts static information of the target file in the installation package of the application to be detected, and processes the static information into a digital feature vector by feature transformation; based on the trained AI model, the digital feature vector is operated by the AI ​​model to obtain the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application. The embodiment of the present invention solves the problems of difficult rule extraction, low coverage, poor scalability, and easy bypass in the traditional malicious application detection based on manually extracted rules, and has higher accuracy and timeliness in malicious program detection.

[0073] Based on the above embodiment, the step 101 of parsing the target file in the installation package of the application to be detected and extracting the static information in the target file specifically includes:

[0074] 101.1, parsing the information of each sub-file in the installation package of the application to be detected to obtain the target file;

[0075] 101.2, extracting a target character string matching a specific keyword from the target file; the target character string includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;

[0076] 101.3, classify and summarize the target character string according to the behavior dimension, the permission dimension and the content dimension to obtain the static information of the application to be detected.

[0077] In step 101.1 of the embodiment of the present invention, a target file is obtained from a sub-file of the installation package of the application to be detected, wherein the target file may be a file such as AndroidManifest.xml, classes.dex, etc. in the sub-file.

[0078] The embodiment of the present invention extracts static information from the target file by keyword matching. The specific keywords described in step 101.2 are keywords related to behavior, authority, and content. The target string matched by the specific keywords is the string information related to behavior, authority, and content. By matching all the specific keywords set in the embodiment of the present invention, all information in the behavior dimension, authority dimension, and content dimension may be matched, or partial information may be matched.

[0079] Step 101.3 categorizes and summarizes the character strings matched by the keywords. All character string information of the behavior dimension is categorized into the behavior dimension, all character string information of the content dimension is categorized into the content dimension, and all character string information matched to the permission dimension is categorized into the permission dimension. After categorization and summary, all information of the behavior dimension, permission dimension and content dimension is static information.

[0080] In an optional embodiment, the static information of the content dimension is numerical data, and the static information of the behavior dimension and the permission dimension is Boolean data. When classifying and aggregating, the character string of the content dimension can be converted into a numerical value for aggregation, and the character string of the behavior dimension and the permission dimension can be converted into a Boolean value for aggregation.

[0081] Specifically, the behavior dimension includes behavior information when the application is running;

[0082] The permission dimension includes permission information required by the application when performing a specific behavior;

[0083] The content dimension includes at least one of the following information: the total file size of the application, the number of files included in the application, the size of specific files in the application, the number of specific components in the application, and application hardening information.

[0084] The application hardening described in the embodiment of the present invention refers to hardening and protecting the application by detecting the running status of the application, intercepting bad behaviors, and preventing malicious programs from exploiting vulnerabilities in the application to damage the computer.

[0085] The static information in the embodiment of the present invention includes at least one dimension information of the behavior dimension, the authority dimension and the content dimension; the specific method for extracting the information of each dimension includes but is not limited to the following description, and each dimension information includes at least one of the descriptions:

[0086] In the behavior dimension, the behavior of an application refers to the behavior performed by the program during operation, such as accessing local files, sending text messages, and connecting to the Internet. Specifically, these behavior information exists in the classes.dex file in the form of specific class names, method names, and variable names. For example, the method name corresponding to the "SMS_Delete" behavior is "api:android / content / ContentResolver->delete,url:content: / / sms;". Based on the analysis of a large number of existing applications, the correspondence between some behaviors used in the embodiment of the present invention and the Android application method names is shown in Table 1.

[0087] Table 1

[0088]

[0089]

[0090] In the permission dimension, when an application reads a file, it needs to declare the permission to read the file first, which includes some sensitive permissions. Therefore, whether certain permissions are declared or not has a certain reference value for judging the maliciousness of the application. Preferably, the embodiment of the present invention is from the "<uses-permission name=”xxx”> " field to obtain information about the permission declaration. Based on the analysis of a large number of existing applications, some permission fields of the Android application used in the embodiment of the present invention are shown in Table 2.

[0091] Table 2

[0092] Permissions android.permission.ACCESS_CHECKIN_PROPERTIES android.permission.ACCESS_COARSE_LOCATION android.permission.ACCESS_FINE_LOCATION android.permission.ACCESS_LOCATION_EXTRA_COMMANDS android.permission.ACCESS_MOCK_LOCATION android.permission.ACCESS_NETWORK_STATE android.permission.ACCESS_SURFACE_FLINGER android.permission.ACCESS_WIFI_STATE android.permission.ACCOUNT_MANAGER android.permission.AUTHENTICATE_ACCOUNTS

[0093] In the content dimension, the total file size of the application, the size of the classes.dex file, the number of included files, and whether the application uses hardening measures are counted; in the AndroidManifest.xml file, the keyword " <activity> ”、" <service> ”、" <receiver>"as well as"<meta_data> "Count the number of activity components, service components, receiver components, and meta_data components respectively.

[0094] The embodiment of the present invention is based on the analysis of a large number of existing applications and extracts application-specific behaviors, permissions, sensitive strings, certificates and other dimensional information contained in the application installation package. Malicious applications are highly distinguishable from normal applications in these dimensions, which are an important basis for judging the maliciousness of applications.

[0095] As mentioned above, the static information extracted by the embodiment of the present invention mainly includes two types of data, one is numerical type data, and the other is Boolean type data. The embodiment of the present invention performs feature transformation on the extracted static information through the following method.

[0096] Based on any of the above optional embodiments, the step 102 of processing the static information into a digital feature vector by feature transformation specifically includes:

[0097] 102.1, based on each dimension data in the static information, perform feature transformation on the numerical data in each dimension data, and convert the feature transformation into a coded number consisting of 0 and 1 through one-hot encoding;

[0098] 102.2, converting the Boolean data in each dimension data into a number 0 or 1 respectively;

[0099] 102.3, concatenate the converted numbers of all numerical data and Boolean data into a digital feature vector composed of 0 and 1 in a preset order.

[0100] Step 102.1 of the embodiment of the present invention performs feature transformation on numerical data, mainly to perform feature transformation on static information of content dimension; based on the content dimension information in the static information, the content dimension information is converted into numerical data, and feature transformation is performed on each numerical data. After the feature transformation, each data in the dimension information will obtain a feature transformation number. For example, assuming that the content dimension contains 4 numerical data, namely 10, 11, 11, and 15, when performing feature transformation, an optional method is to classify 10, 11, and 11 into one category and give a feature transformation number 1, and to classify 15 as another category and give a feature transformation number 2, so that the feature transformation numbers of numerical data 10, 11, 11, and 15 after feature transformation are 1, 1, 1, and 2 respectively, so that the data can be discretized to improve the adaptability and generalization ability of the AI ​​model. The above examples are only hypothetical optional solutions, and the embodiments of the present invention do not limit other solutions based on specific application requirements.

[0101] The feature transformation number is then one-hot encoded, that is, converted into a coded number consisting of 0 and 1. One-hot encoding is also called one-bit effective encoding. Its method is to use an N-bit state register to encode N states. Each state has its own independent register bit, and at any time, only one of them is valid.

[0102] For example, to encode six states:

[0103] The natural sequence code is: 000, 001, 010, 011, 100, 101;

[0104] The one-hot encoding is: 000001, 000010, 000100, 001000, 010000, 10000.

[0105] It should be noted that the embodiment of the present invention mainly describes step 102.1 based on information of content dimension, and does not exclude the possibility that numerical data of other dimensions may be transformed in features through step 102.1.

[0106] Step 102.2 performs feature transformation on Boolean data, mainly on static information of the behavior dimension and the permission dimension (not excluding other dimensions). If the string of the application can be matched through the specific keyword, it is 1, otherwise it is 0. The embodiment of the present invention does not limit the order of implementation of steps 102.1 and 102.2.

[0107] Step 102.3 concatenates all the coded numbers converted from the numerical data and the numbers 0 and 1 converted from the Boolean data into a character string in a preset order. The embodiment of the present invention does not limit the concatenation order as long as it is consistent with the concatenation order during AI model training.

[0108] The feature transformation method of step 102.1 in the embodiment of the present invention is flexible and diverse. Different feature transformation methods can be adopted based on the specific data of the extracted static information and the computing requirements of different scenarios.

[0109] Based on any of the above optional embodiments, step 102.1 converts the numerical data in each dimension data into a coded number consisting of 0 and 1 through one-hot encoding after feature transformation, specifically including:

[0110] Based on the numerical data in each dimension data, obtain N numerical segments according to a first preset rule, and sort the N numerical segments according to numerical values ​​to obtain a ranking of each numerical segment, where N is an integer greater than 0;

[0111] Matching each numerical data in each dimension data with the N numerical segments according to a second preset rule, so that each numerical data matches one numerical segment, and using the ranking of the matched numerical segments as the characteristic transformation number of each numerical data;

[0112] The feature transformation digits of each numerical data are uniquely encoded to obtain encoded digits consisting of 0 and 1.

[0113] Figure 2 The figure is a numerical segmentation diagram of an embodiment of the present invention. The embodiment of the present invention is segmented according to the numerical data of each dimension; specifically, the numerical data of the embodiment of the present invention is mainly data of the content dimension, and the content dimension will extract multiple different content information, such as the number of components A, the number of components B, etc., so the content dimension also includes multiple dimensions, such as the number of components A dimension, the number of components B dimension, etc.

[0114] Preferably, each dimension data described in the embodiment of the present invention may refer to a dimension in a dimension. Taking the number of A components as an example, assuming that the data about the number of A components extracted by the embodiment of the present invention is 1150 in total, that is, 1150 numerical data are obtained, then segmentation is performed according to the total 1150 numerical data, and the segmentation method is the first preset rule; the first preset rule may be segmentation according to the numerical size of the 1150 numerical data, or segmentation according to the number of data of the 1150 numerical data, or segmentation according to a mixture of numerical size and number of data. The specific segmentation method may be determined according to demand, and the embodiment of the present invention does not limit this. Similarly, the number of B components for the content dimension is also segmented according to the total number of B components extracted; the segmentation methods for other information extraction are similar.

[0115] like Figure 2 As shown, it is assumed that the embodiment of the present invention segments the 1150 numerical data in the example through the first preset rule to obtain N numerical segments. After the N numerical segments are sorted by size, the sorting positions are 1, 2, ..., N respectively.

[0116] After the numerical value segments are sorted, the embodiment of the present invention "matches" all the numerical data according to the second preset rule, wherein the second preset rule is a matching method corresponding to the first preset rule, that is: if the first preset rule is to segment by numerical value size, then the second preset rule matches by numerical value size; if the first preset rule is to segment by the number of data, then the second preset rule matches by the number of data; if the first preset rule is to segment by a mixture of numerical value size and data number, then the second preset rule matches by a mixture of numerical value size and data number.

[0117] Please refer to Figure 2 , taking the first preset rule of segmenting by numerical value as an example, for the 1150 numerical data in the above example, the 1150 numerical data are matched with N numerical segments respectively. If the values ​​of 50 of them are within the numerical range of segment 1, the ranking rank 1 is used as the feature transformation number of these 50 data, that is, each of the 50 data obtains the feature transformation number 1; if the values ​​of 100 of them are within the numerical range of segment 2, the ranking rank 2 is used as the feature transformation number of these 100 data, that is, each of the 100 data obtains the feature transformation number 2; and so on, each numerical data will obtain a feature transformation number, which will not be repeated.

[0118] After segmented matching of each numerical data, each numerical data obtains a feature transformation number. The embodiment of the present invention performs one-hot encoding on each feature transformation number to obtain a coded number composed of 0 and 1; the one-hot encoding process of the following embodiment is the same as this and will not be repeated hereafter.

[0119] The embodiment of the present invention performs feature transformation on numerical data, and can classify and discretize numerical data with similar characteristics (for example, the numerical values ​​are different but very close), which can improve the generalization ability of the AI ​​model during AI model training. Preferably, the embodiment of the present invention provides three feature transformation methods, namely, feature transformation methods by numerical value, by number of data, and by a mixture of numerical value and number of data, as follows:

[0120] The first feature transformation method: the first preset rule is to segment by value, and the second preset rule is to match by value;

[0121] The second feature transformation method: the first preset rule is to segment by the number of data, and the second preset rule is to match by the number of data;

[0122] The third feature transformation method: the first preset rule is to segment by a mixture of numerical value size and data number, and the second preset rule is to match by a mixture of numerical value size and data number.

[0123] Each feature transformation method is described in detail below.

[0124] Based on any of the above optional embodiments, step 102.1, after performing feature transformation on the numerical data in each dimension data, converting the numerical data into a coded number consisting of 0 and 1 through one-hot encoding, specifically includes:

[0125] Based on the numerical data in each dimension data, the maximum value and the minimum value in the numerical data are used as the numerical interval, the numerical interval is divided into N parts to obtain N numerical segments, and the N numerical segments are sorted according to the numerical size to obtain the ranking of each numerical segment, where N is an integer greater than 0;

[0126] For each numerical data in each dimension data, if the value of the numerical data is within the numerical range of a numerical segment S, the ranking of the numerical segment S is used as the characteristic transformation number of the numerical data;

[0127] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0128] This embodiment is the first feature transformation method, which is suitable for scenarios where the values ​​of numerical data are continuously distributed. The following example is used to illustrate. Assume that the 1150 numerical data in the above example are continuously distributed, with a maximum value of 1000 and a minimum value of 0; take N = 10, the numerical interval is 1000, and the numerical interval 1000 is evenly divided into 10 segments, and the numerical range of each segment is [0, 100], [101, 200], [201, 300], [301, 400], [401, 500], [501, 600], [601, 700], [701, 800], [801, 900], [901, 1000], and the numerical intervals are sorted by the size of the numerical range. The sorting positions corresponding to the above numerical segments are 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 respectively.

[0129] Match the 1150 numerical data with the above 10 numerical segments respectively. If the value of a numerical data a is 10 and is in the [0, 100] segment, then the ranking 1 of the [0, 100] segment is used as the feature transformation number of a; and so on, perform segment matching on all numerical data respectively, then each numerical data will obtain a feature transformation number ranging from 1 to 10. Thus, the 1150 numerical data are divided into 10 parts, one part obtains the feature transformation number 1, one part obtains the feature transformation number 2, and so on, one part obtains the feature transformation number 10.

[0130] It should be noted that the N numerical segments can be sorted in ascending order or in descending order, and the sorting method during malicious application detection can be consistent with the sorting method during AI model training.

[0131] In addition, for the value at the boundary point of the numerical segment, it can be matched to a smaller numerical segment or a larger numerical segment according to the agreed rules, and the embodiment of the present invention is not limited to this.

[0132] In addition, the above example divides the numerical interval 1000 into 10 segments evenly. If a specific rule is used to perform non-even segmentation, it also falls within the protection scope of the embodiment of the present invention.

[0133] For the numerical data that have obtained the feature transformation number, one-hot encoding is performed. For example, if the feature transformation number of the numerical data a is 1 and the total feature transformation numbers are 10, then the one-hot encoding is 1000000000; if the feature transformation number of the numerical data b is 2, then the one-hot encoding is 0100000000; and so on.

[0134] Based on any of the above optional embodiments, step 102.1, after performing feature transformation on the numerical data in each dimension data, converting the numerical data into a coded number consisting of 0 and 1 through one-hot encoding, specifically includes:

[0135] Based on the numerical data in each dimension data, segment the data according to the total number of the numerical data to obtain N numerical segments, and sort the N numerical segments according to the numerical values ​​to obtain the ranking of each numerical segment, where N is an integer greater than 0;

[0136] Sort the numerical data in each dimension data by size, and obtain the sorting number of each numerical data;

[0137] For each numerical data, the ranking of the numerical segment corresponding to the ranking number of each numerical data is used as the characteristic transformation number of each numerical data;

[0138] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0139] This embodiment is the second feature transformation method, which is suitable for scenarios where the values ​​of numerical data are discretely distributed. The following example is used to illustrate. Assuming that the 1150 numerical data in the above example have discrete distribution values, if the first feature transformation method is used, some numerical segments may have no matching data, and some numerical segments may match a lot of data, which will have an adverse effect on the data calculation of the AI ​​model.

[0140] In order to avoid the above situation, the second feature transformation method uses the number of data for segmentation and matching. Assuming the 1150 numerical data in the above example, take N=10, then 10 numerical segments are obtained, and each numerical segment can match 115 numerical data; sort the 1150 numerical data by numerical value, then the 1st to 115th numerical data in the sorting match numerical segment 1, and obtain feature transformation number 1, the 116th to 230th numerical data match numerical segment 2, and obtain feature transformation number 2, and so on, the 1036th to 1150th numerical data match numerical segment 10, and obtain feature transformation number 10. The unique hot encoding of the numerical data that has obtained the feature transformation number is the same as the above embodiment, and will not be repeated here.

[0141] It should be noted that the above example divides the number of data 1150 into 10 segments evenly. If a specific rule is used to perform non-even segmentation, it also falls within the protection scope of the embodiment of the present invention.

[0142] Based on any of the above optional embodiments, step 102.1, after performing feature transformation on the numerical data in each dimension data, converting the numerical data into a coded number consisting of 0 and 1 through one-hot encoding, specifically includes:

[0143] Based on the numerical data in each dimension data, segment part of the numerical data according to the number of data to obtain n numerical segments, where n is an integer greater than 0;

[0144] Eliminate some of the numerical data that are segmented according to the number of data, and use the maximum and minimum values ​​of the remaining data as the numerical interval, divide the numerical interval into Nn parts, and obtain Nn numerical segments, where N is an integer greater than n;

[0145] Sort the N numerical segments according to their numerical values ​​to obtain the ranking of each numerical segment;

[0146] Sort the numerical data in each dimension data by size, match them with N numerical segments respectively, and use the ranking of the matched numerical segments as the feature transformation number of the numerical data;

[0147] Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

[0148] This embodiment is the third feature transformation method, which is suitable for scenarios where the values ​​of numerical data have both discrete distribution and continuous distribution, and is explained below by example. It should be noted that the order of text description of the third feature transformation method mentioned above is not used to limit the order of steps in actual implementation; in fact, all methods of segmenting by mixing numerical value size and data number, and matching by mixing numerical value size and data number, all fall within the protection scope of the embodiments of the present invention.

[0149] Assuming the 1150 numerical data in the above example, the 1150 numerical data are sorted by numerical value, wherein the first segment of data and the last segment of data are relatively discrete, and the data in the middle are relatively continuous; then the first segment of data, i.e., x% of the total number of data, is taken as the first numerical segment, and the sorting position is 1, and the last segment of data, i.e., y% of the total number of data, is taken as the Nth numerical segment, and the sorting position is N. Take N=10, x=10, y=20, 1150*10%=115, 1150*20%=230, then the first numerical segment can match 115 numerical data, and the Nth numerical segment can match 230 numerical data; after sorting the 1150 numerical data by numerical value, the 1st to 115th numerical data are matched to the first numerical segment, and the feature transformation number 1 is obtained, and the 921st to 1150th numerical data are matched to the Nth numerical segment, and the feature transformation number 10 is obtained.

[0150] The maximum and minimum values ​​of the 116th to 920th numerical data among the 1150 numerical data are taken as the numerical interval, and the numerical interval is divided into 10-2 segments, i.e., 8 numerical segments, by an average or non-average method, and the ranking positions are 2, 3, 4, 5, 6, 7, 8, and 9 respectively; each numerical data from the 116th to 920th numerical data is matched with numerical segments 2 to 9 respectively, and if the value of the numerical data is within the numerical range of a numerical segment S, the ranking position of the numerical segment S is taken as its characteristic transformation number.

[0151] The one-hot encoding of the numerical data for which the feature transformation numbers have been obtained is the same as in the above embodiment and will not be repeated here.

[0152] It should be noted that, in the above example, the first and last segments of the sorted numerical data are segmented according to the number of data, and the middle part is segmented according to the numerical size range. It is also possible to segment the first and last segments according to the numerical size range and the middle part according to the number of data according to the needs of time application, and so on. All data segmentation that includes segmentation according to the numerical size range and segmentation according to the number of data falls within the protection scope of the embodiments of the present invention.

[0153] The embodiment of the present invention performs feature transformation on the extracted static information and then performs one-hot encoding based on the data itself. When training the AI ​​model, it ensures that the information contained in the data can be extracted to the maximum extent while ensuring the final effect of the AI ​​model; when detecting malicious applications, it ensures the accuracy of detection to the greatest extent.

[0154] Based on any of the above optional embodiments, the AI ​​model network structure of the embodiment of the present invention includes but is not limited to convolutional neural networks, recurrent neural networks, deep learning networks, machine learning networks, etc., and mainly includes four parts: input layer, decomposition layer, hidden layer and output layer. Figure 3 Schematic diagram of the network structure of the AI ​​model according to the embodiment of the present invention. Figure 3 As shown, the network structure of the AI ​​model includes four parts: input layers, factorization machines, hidden layers and output layers;

[0155] The input layer is used to receive a digital feature vector and a malicious application label as input data; the decomposition layer is used to extract low-order features from the input data and perform calculations based on the low-order features; the hidden layer is used to extract high-order features from the input data, calculate the malicious features of the application based on the high-order features, and segment the malicious features of applications of different maliciousness from a high-dimensional space; the output layer is used to merge the calculation results of the decomposition layer and the calculation results of the hidden layer, and output the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0156] Figure 4 A schematic diagram of the AI ​​model training process according to an embodiment of the present invention. Figure 4 As shown, the AI ​​model described in the embodiment of the present invention is trained by the following method:

[0157] 400, obtaining a sample of an application installation package and a malicious label of the sample of the application installation package;

[0158] 401, parsing a second target file in the sample application installation package, and extracting second static information in the second target file; the second static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension;

[0159] 402, processing the second static information into a second digital feature vector by feature transformation, where the second digital feature vector consists of 0 and 1;

[0160] 403. Input the second digital feature vector and the malicious label of the application installation package sample into the constructed AI model, and train the AI ​​model to obtain an AI model that meets the expected requirements.

[0161] It should be noted that the second static information and the second digital feature vector in the embodiment of the present invention are only used to make a noun distinction from the static information and the digital feature vector mentioned above, where "second" has no practical meaning; the practical meanings of the second static information and the static information are exactly the same, and the practical meanings of the second digital feature vector and the digital feature vector are exactly the same.

[0162] In the embodiment of the present invention, the application installation package sample data and malicious label data obtained in step 400 can be the result of manual analysis, and the sample collection and maliciousness judgment are performed manually. In the embodiment of the present invention, a large number of complete Android platform applications are collected, and the characteristics of malicious Trojans and other applications are manually analyzed to extract the rules for judging the maliciousness of the application. Then, the maliciousness of the collected Android platform applications is judged based on the rules to obtain the malicious labels corresponding to the application installation package samples.

[0163] As mentioned above, the AI ​​model training phase and the application malicious detection phase of the embodiment of the present invention both include static information extraction and feature transformation of the extracted static information, and their processing methods are exactly the same.

[0164] Step 401 of the AI ​​model training method of the embodiment of the present invention is exactly the same as step 101 of the AI ​​model-based malicious application detection method and the processing methods of all optional embodiments of step 101; if there are multiple different optional embodiments, the effect of malicious application detection can be achieved by keeping the optional embodiments adopted in steps 101 and 401 consistent, which will not be repeated here.

[0165] Step 402 of the AI ​​model training method of the embodiment of the present invention is exactly the same as step 102 of the AI ​​model-based malicious application detection method and the processing methods of all optional embodiments of step 102; if there are multiple different optional embodiments, the effect of malicious application detection can be achieved by keeping the optional embodiments adopted in steps 102 and 402 consistent, which will not be repeated here.

[0166] Step 403 of the AI ​​model training method in the embodiment of the present invention is different from step 103 of the malicious application detection method based on the AI ​​model in that the input data of step 403 includes not only the digital feature vector after feature transformation, but also the malicious label of the application installation package sample. The AI ​​model is trained by the digital feature vector and the malicious label of the application installation package sample, so that the AI ​​model has the ability to identify malicious applications based on the digital feature vector.

[0167] The data received by the input layer (Input Layers) of the AI ​​model of the embodiment of the present invention are digital feature vectors formed after feature transformation. These digital feature vectors contain features extracted from dimensions such as permissions, behaviors, and content to characterize the maliciousness of the application APP. The numbers in the digital feature vectors are composed of 0 and 1.

[0168] The factorization machine layer (Factorization Machines) of the AI ​​model of the embodiment of the present invention is used to extract low-order features from the data. The maliciousness of an APP is often a series of actions. Only by coordinating behaviors, permissions, content, etc. at the same time can the malicious behavior be completed. Therefore, the combination of features can often reflect the maliciousness of the APP to a large extent. The factorization machine mainly extracts feature combinations through the inner product of the hidden variables of each dimensional feature to ensure that the model can save the combination information between features.

[0169] Preferably, the hidden layer (Hidden Layers) of the AI ​​model in the embodiment of the present invention is a feedforward neural network, including hidden units (Hidden Unites) and a Sigmoid function (Sigmoid Function); the hidden layer includes 6 hidden layers, and the number of neuron nodes in each layer ranges from 32 to 256. The specific number of neuron nodes in each layer can be 64, 128, 256, 128, 64 or 32, which is used to extract high-order features in the data and segment the features of APPs with different maliciousness from a high-dimensional space.

[0170] The output layer (Output Layers) of the AI ​​model of the embodiment of the present invention merges the forward calculation results of the decomposition machine layer and the hidden layer, and the number of nodes outputted finally is 2 or 1: if it is 2, it is the probability that this sample is a non-malicious application and the probability that it is a malicious application respectively; if it is 1, it is the probability that this sample is a non-malicious application or the probability that it is a malicious application.

[0171] Preferably, by setting probability thresholds for malicious applications and non-malicious applications, detection requirements for low false alarm scenarios can be met.

[0172] Finally, the model is trained through back propagation; using cross-validation as a means, with accuracy, recall, precision, and F-value as standards, and taking into account the influence of time factors and business background, the parameters, number of iterations, model structure, etc. of the AI ​​model are adjusted to finally obtain an AI model that meets the expected requirements.

[0173] Based on any of the above optional embodiments, step 403 inputs the second digital feature vector and the malicious label of the application installation package sample into the built AI model to train the AI ​​model, specifically including:

[0174] 403.1, converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector;

[0175] 403.2. Input the second digital feature vector and the malicious label of the second digital feature vector into the constructed AI model to train the AI ​​model.

[0176] In the embodiment of the present invention, the digital feature vector is obtained by performing feature transformation on static information extracted from the application installation package sample, and the malicious label of the application installation package sample is the result of maliciously labeling the APP after analyzing a large number of mainstream malware on the Android platform. One application installation package sample corresponds to one digital feature vector and one malicious label at the same time. Therefore, the digital feature vector of the application installation package sample can be matched with the malicious label, and the malicious label of the application installation package sample can be converted into a malicious label of the digital feature vector, that is, one digital feature vector corresponds to one malicious label.

[0177] In this way, one input data of the AI ​​model of the embodiment of the present invention is a digital feature vector, and the other input data is a malicious label corresponding to the digital feature vector, and the AI ​​model is trained.

[0178] The embodiment of the present invention matches the digital feature vector with the malicious label. By labeling the digital feature vector, the generalization ability of the AI ​​model can be improved.

[0179] Based on any of the above optional embodiments, the step 403.1 of converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes:

[0180] Based on the malicious labels of all application installation package samples, taking the malicious label of any application installation package sample as the malicious label of the second digital feature vector corresponding to the any application installation package sample;

[0181] Based on all the second digital feature vectors and their corresponding malicious labels, data in which the second digital feature vectors and their corresponding malicious labels are completely identical are deduplicated to obtain the deduplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

[0182] The embodiment of the present invention converts the malicious label of the application installation package sample into the malicious label of the digital feature vector, and it is inevitable that there are data in which the digital feature vector and the malicious label are exactly the same. For example, application A is malicious after analysis, application B is malicious after analysis, and application C is malicious after analysis; assuming that after application A, B and C extract static information for feature transformation, the obtained digital feature vectors are exactly the same, assuming 0010000000, and the malicious labels of A, B and C are all malicious, assuming 1, then three completely identical data will appear (0010000000, 1). The embodiment of the present invention deduplicates the data with exactly the same digital feature vector and its corresponding malicious label, which can greatly reduce the amount of data, thereby improving the generalization ability of the AI ​​model and accelerating the AI ​​model training process. After laboratory scenario testing, after data deduplication, the amount of data is reduced by about 1 order of magnitude.

[0183] Furthermore, if there are multiple identical digital feature vectors, the corresponding malicious labels include malicious and non-malicious. If the number of malicious ones is greater than the number of non-malicious ones, the digital feature vector is marked as malicious, otherwise it is marked as non-malicious.

[0184] For example, after extracting static information and performing feature transformation on applications A, B, C, D, and E, the resulting digital feature vectors are exactly the same, assuming they are 0010000000, and the malicious labels of A, B, and C are all malicious, assuming they are 1, and the malicious labels of D and E are both non-malicious, assuming they are 0. Then, the number of digital feature vectors 0010000000 corresponding to malicious labels of 1 is greater than the number of corresponding malicious labels of 0, so the digital feature vector 0010000000 is marked as 1.

[0185] This embodiment deduplicates and relabels the digital feature vectors according to the number of malicious and non-malicious labels corresponding to the digital feature vectors, further reducing the amount of data, alleviating the data processing pressure of the AI ​​model, and also further improving the generalization ability of the AI ​​model.

[0186] The training data of the AI ​​model of the embodiment of the present invention reaches millions of levels. In a training environment, the number of malicious applications reaches more than 3.96 million, and the number of non-malicious applications reaches more than 880,000, that is, the total training data is about 5 million. The digital feature vectors of all applications are obtained, and the corresponding digital feature vectors are labeled and deduplicated using the malicious labels of the applications, and about 500,000 digital feature vectors with malicious or non-malicious labels are obtained; about 220,000 digital feature vectors are labeled as malicious, and about 280,000 digital feature vectors are labeled as non-malicious. A training set is selected from the above data (labeled and deduplicated digital feature vectors) to train the AI ​​model.

[0187] During the test process, a strategy of not overlapping with the training set was adopted to select test data. Specifically, 121,946 malicious applications and 81,757 non-malicious applications were selected. The test results are shown in Table 3.

[0188] As can be seen from Table 3, among the total 121,946 training data, 119,508 data are actually malicious and predicted to be malicious; 81,349 data are actually non-malicious and predicted to be non-malicious; 2,348 data are actually malicious and predicted to be non-malicious; and 408 data are actually non-malicious and predicted to be malicious.

[0189] Table 3

[0190]

[0191]

[0192] Thus, the various indicators of the AI ​​model of the embodiment of the present invention can be calculated as follows:

[0193] The detection rate is: 119508 / (119508+2348)=98.00%;

[0194] The false alarm rate is: 408 / (408+81349)=0.50%;

[0195] The accuracy rate is: (119508+81349) / 203703=98.60%;

[0196] It can be seen that the AI ​​model of the embodiment of the present invention has extremely high detection rate and accuracy, extremely low false alarm rate, and has very good beneficial effects.

[0197] In summary, the embodiment of the present invention firstly extracts several features such as file size, whether it is packed, the number of supported languages, etc. from the three dimensions of behavior, permission and content based on the analysis of Android platform applications and combined with mobile application security knowledge; secondly, the feature data is preprocessed by feature transformation; thirdly, multiple rounds of model training are carried out using a model redesigned based on a deep learning algorithm, and the model effect is evaluated to achieve the expected goal; finally, features are extracted from applications with unknown maliciousness, and predictions are made through the model to complete the maliciousness determination of the application.

[0198] Among them, static information is the cornerstone of AI model training and the basis for malicious application detection. Static information contains application-specific behaviors, permissions, sensitive strings, certificates and other information. Malicious applications are highly distinguishable from normal applications in these specific dimensions, which is an important basis for judging the maliciousness of applications.

[0199] Feature transformation is the guarantee of AI model effectiveness. Feature transformation is based on the data itself. From the perspective of the model, it ensures that the information contained in the data can be extracted to the maximum extent possible while ensuring the final effect of the model.

[0200] Model design is the key to the success of AI models. A properly designed model can maximize the conversion of information in the data into knowledge, and then solidify the knowledge for judging the maliciousness of applications.

[0201] In summary, the embodiment of the present invention extracts static features and statistical features of data from a large number of Android platform applications, redesigns the algorithm based on a deep learning algorithm, and then performs model training on the data after feature extraction, and finally obtains an AI model for detecting malicious applications on the Android platform. The AI ​​model is used to detect malicious applications, which is independent of the update of the malicious feature library and has strong learning ability, providing a new direction for malicious detection of Android platform applications. The present invention solves the problems of difficult rule extraction, low coverage, poor scalability, and easy bypassing in traditional malicious application detection based on manually extracted rules, and the embodiment of the present invention has higher accuracy and timeliness.

[0202] Figure 5 The embodiment of the present invention is a malicious application detection device based on an AI model. The embodiment of the present invention also provides a malicious application detection device based on an AI model, such as Figure 5 As shown, it includes a static information extraction module 501, a feature transformation module 502 and a malicious application detection module 503:

[0203] The static information extraction module 501 is used to parse the target file in the installation package of the application to be detected, and extract the static information in the target file, wherein the static information includes at least one dimension information of the behavior dimension, the permission dimension and the content dimension;

[0204] The feature transformation module 502 is used to process the static information into a digital feature vector by feature transformation, and the digital feature vector is composed of 0 and 1;

[0205] The malicious application detection module 503 is used to input the digital feature vector into a trained AI model to obtain a malicious detection result for the application to be detected; the AI ​​model is used to perform operations based on the input digital feature vector and output the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that it is a non-malicious application.

[0206] The malicious application detection device based on the AI ​​model in the embodiment of the present invention is used to execute Figure 1 The technical solution of the AI ​​model-based malicious application detection method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.

[0207] Figure 6 This is a schematic diagram of the electronic device framework according to an embodiment of the present invention. Figure 6 , an embodiment of the present invention provides an electronic device, including: a processor (processor) 610, a communication interface (Communications Interface) 620, a memory (memory) 630 and a bus 640, wherein the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the bus 640. The processor 610 can call the logic instructions in the memory 630 to execute the following method, including: parsing the target file in the installation package of the application to be detected, extracting the static information in the target file, the static information includes at least one dimension information of the behavior dimension, the permission dimension and the content dimension; processing the static information into a digital feature vector by feature transformation, the digital feature vector is composed of 0 and 1; inputting the digital feature vector into a trained AI model to obtain the maliciousness detection result of the application to be detected; the AI ​​model is trained according to the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application.

[0208] An embodiment of the present invention discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments, for example, including: parsing a target file in an installation package of an application to be detected, extracting static information in the target file, and the static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; processing the static information into a digital feature vector by a feature transformation method, and the digital feature vector is composed of 0 and 1; inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result for the application to be detected; the AI ​​model is trained according to the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application.

[0209] An embodiment of the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the methods provided by the above-mentioned method embodiments, for example, including: parsing a target file in an installation package of an application to be detected, extracting static information in the target file, and the static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; processing the static information into a digital feature vector through feature transformation, and the digital feature vector consists of 0 and 1; inputting the digital feature vector into a trained AI model to obtain a malicious detection result for the application to be detected; the AI ​​model is trained according to the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability of a non-malicious application.

[0210] A person skilled in the art can understand that the implementation of the above-mentioned device embodiment or method embodiment is only illustrative, wherein the processor and the memory may be physically separated components or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person skilled in the art can understand and implement it without paying creative labor.

[0211] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a USB flash drive, a mobile hard disk, ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / receiver> < / service> < / activity>

Claims

1. A malicious application detection method based on AI model, It is characterized in that include: Parsing a target file in an installation package of the application to be detected, and extracting static information in the target file, wherein the static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; Processing the static information into a digital feature vector by feature transformation, wherein the digital feature vector is composed of 0 and 1; Inputting the digital feature vector into a trained AI model to obtain a maliciousness detection result of the application to be detected; the AI ​​model is trained based on the input digital feature vector, and outputs the probability that the application corresponding to the digital feature vector is a malicious application and / or the probability that the application is a non-malicious application; The behavior dimension includes behavior information when the application is running; The permission dimension includes permission information required by the application when performing a specific behavior; The content dimension includes at least one of the following information: the total file size of the application, the number of files included in the application, the size of specific files in the application, the number of specific components in the application, and application hardening information; the processing of the static information into a digital feature vector by feature transformation specifically includes: Based on each dimension data in the static information, the numerical data in each dimension data is subjected to feature transformation, and then converted into a coded number consisting of 0 and 1 through one-hot encoding; Convert the Boolean data in each dimension data into a number 0 or 1 respectively; All the converted numbers of numeric data and Boolean data are concatenated into a digital feature vector consisting of 0 and 1 in a preset order; After the numerical data in each dimension data is subjected to feature transformation, it is converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including: Based on the numerical data in each dimension data, obtain N numerical segments according to a first preset rule, and sort the N numerical segments according to numerical values ​​to obtain a ranking of each numerical segment, where N is an integer greater than 0; Matching each numerical data in each dimension data with the N numerical segments according to a second preset rule, so that each numerical data matches one numerical segment, and using the ranking of the matched numerical segments as the characteristic transformation number of each numerical data; The feature transformation digits of each numerical data are uniquely encoded to obtain encoded digits consisting of 0 and 1.

2. The method according to claim 1, It is characterized in that After the numerical data in each dimension data is subjected to feature transformation, it is converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including: Based on the numerical data in each dimension data, the maximum value and the minimum value in the numerical data are used as the numerical interval, the numerical interval is divided into N parts to obtain N numerical segments, and the N numerical segments are sorted according to the numerical size to obtain the ranking of each numerical segment, where N is an integer greater than 0; For each numerical data in each dimension data, if the value of the numerical data is within the numerical range of a numerical segment S, the ranking of the numerical segment S is used as the characteristic transformation number of the numerical data; Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

3. The method according to claim 1, It is characterized in that After the numerical data in each dimension data is subjected to feature transformation, it is converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including: Based on the numerical data in each dimension data, segment the data according to the total number of the numerical data to obtain N numerical segments, and sort the N numerical segments according to the numerical values ​​to obtain the ranking of each numerical segment, where N is an integer greater than 0; Sort the numerical data in each dimension data by size, and obtain the sorting number of each numerical data; For each numerical data, the ranking of the numerical segment corresponding to the ranking number of each numerical data is used as the characteristic transformation number of each numerical data; Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

4. The method according to claim 1, It is characterized in that After the numerical data in each dimension data is subjected to feature transformation, it is converted into a coded number consisting of 0 and 1 through one-hot encoding, specifically including: Based on the numerical data in each dimension data, segment part of the numerical data according to the number of data to obtain n numerical segments, where n is an integer greater than 0; Eliminate some of the numerical data that are segmented according to the number of data, and use the maximum and minimum values ​​of the remaining data as the numerical interval, divide the numerical interval into Nn parts, and obtain Nn numerical segments, where N is an integer greater than n; Sort the N numerical segments according to their numerical values ​​to obtain the ranking of each numerical segment; Sort the numerical data in each dimension data by size, match them with N numerical segments respectively, and use the ranking of the matched numerical segments as the feature transformation number of the numerical data; Perform unique hot encoding on the feature transformation number of each numerical data to obtain a coded number consisting of 0 and 1.

5. The method according to any one of claims 1 to 4, It is characterized in that The network structure of the AI ​​model includes four parts: input layer, decomposition layer, hidden layer and output layer; the AI ​​model is trained by the following method: Obtaining application installation package samples and malicious labels of the application installation package samples; Parsing a second target file in the application installation package sample, and extracting second static information in the second target file; the second static information includes at least one dimension information of a behavior dimension, a permission dimension, and a content dimension; Processing the second static information into a second digital feature vector by feature transformation, wherein the second digital feature vector consists of 0 and 1; The second digital feature vector and the malicious label of the application installation package sample are input into the constructed AI model, and the AI ​​model is trained to obtain an AI model that meets the expected requirements.

6. The method according to claim 5, It is characterized in that The inputting the second digital feature vector and the malicious label of the application installation package sample into the built AI model to train the AI ​​model specifically includes: Converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector; The second digital feature vector and the malicious label of the second digital feature vector are input into the constructed AI model to train the AI ​​model.

7. The method according to claim 6, It is characterized in that The converting the malicious label of the application installation package sample into the malicious label of the second digital feature vector specifically includes: Based on the malicious labels of all application installation package samples, taking the malicious label of any application installation package sample as the malicious label of the second digital feature vector corresponding to the any application installation package sample; Based on all the second digital feature vectors and their corresponding malicious labels, data in which the second digital feature vectors and their corresponding malicious labels are completely identical are deduplicated to obtain the deduplicated second digital feature vectors and the malicious labels of the second digital feature vectors.

8. An electronic device, It is characterized in that include: at least one processor; as well as at least one memory in communication with the processor, wherein: The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 7 by calling the program instructions.

9. A non-transitory computer-readable storage medium, It is characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Word vector- and depth neural network-based Android malicious code detection method

    CN108959924A

  • Android malicious software detection method and system based on RNN and CNN

    CN110489968A