Malicious application detection method and system, equipment, medium and product
By obtaining the feature information of the target application and building a feature matrix, combining multiple base classifiers and metaclassifiers, optimizing the parameters of the detection model, the problem of inaccurate detection of malicious applications in the existing technology is solved, and the accuracy and reliability of the detection model are improved.
Patent Information
- Application Number
- CN202411822280.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, malicious application detection is inaccurate, mainly due to the problem of category imbalance, the detection model tends to be mostly classes (benevolent software), while a few classes (malware) samples are easily misclassified.
By obtaining the sensitive permission characteristics, system call characteristics and API call characteristics of the target application, the target feature matrix is constructed, and base classifiers such as random forest RF, support vector machine SVM, logistic regression LR and K proximity algorithm KNN are used, combined with meta-classifiers, the parameters of the detection model are optimized to improve the accuracy of judgment of the unbalanced data set.
It improves the accuracy and reliability of the malicious application detection model, reduces the impact of classification errors of a single classifier, enhances the ability to identify a few classes, and improves the accuracy of judging imbalanced data sets.
Smart Images

Figure CN119939580A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of malicious application detection, and in particular to a method, system, device, medium, and product for detecting malicious applications. Background Art
[0002] In the era of mobile internet, Android smartphones are ubiquitous. However, due to imperfect market regulation, malware has surged, posing a serious threat to user privacy, property, and even personal safety. Given the current severe security situation, malware detection technology for the Android platform is particularly important and urgent. Related technologies have proposed combining machine learning algorithms to detect Android malware. In classification tasks, machine learning algorithms rely on labeled sample datasets to train prediction models and achieve automated predictions based on historical data. However, the class balance of training sample datasets significantly impacts model performance. In real-world scenarios, malware only accounts for 8%-12% of the sample dataset. Class imbalance in datasets can cause detection models to favor the majority class (benign software) while easily misclassifying samples from the minority class (malware). Therefore, addressing class imbalance is crucial for improving the accuracy and reliability of Android malware detection models. Summary of the Invention
[0003] Based on this, a method, system, device, medium and product for detecting malicious applications are provided to solve the problem of inaccurate malicious application detection in the prior art.
[0004] In one aspect, a method for detecting malicious applications is provided, the method comprising:
[0005] Obtaining a target feature information set of a target application; the target feature information set includes at least one of the following: sensitive permission features, system call features, and API call features of the target application;
[0006] Determining a target feature matrix for a target application based on the target feature information set;
[0007] The target feature matrix is input into a target model, so that a base classifier in the target model outputs a detection probability of the target application according to the target feature matrix, and a meta-classifier outputs detection information of the target application according to the detection probability and the target feature matrix; wherein the detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application; the base classifier includes at least two of the following: random forest RF, support vector machine SVM, logistic regression LR, and K-nearest neighbor algorithm KNN.
[0008] Optionally, obtaining a target feature information set of a target application includes:
[0009] Performing static analysis on the installation package of the target application to obtain the sensitive permission characteristics and the API call characteristics;
[0010] The target application is dynamically run to obtain the system call characteristics.
[0011] Optionally, it is characterized in that the parameter set of the base classifier is obtained by the following steps:
[0012] Determining an initial parameter set of the base classifier;
[0013] Substituting the multiple initial parameter sets into the to-be-trained model to obtain multiple objective function values corresponding to the multiple initial parameter sets;
[0014] Determining a predicted mean and variance based on the multiple initial parameter sets and the objective function value;
[0015] Based on the predicted mean and variance, the next initial parameter set of the base classifier is determined until a preset iteration stop condition is met, thereby obtaining the parameter set of the base classifier of the target model.
[0016] Optionally, substituting the multiple initial parameter sets into the to-be-trained model to obtain multiple objective function values corresponding to the multiple initial parameter sets includes:
[0017] Substituting the multiple initial parameter sets into the to-be-trained model respectively;
[0018] Corresponding to each initial parameter set, inputting a preset feature matrix into the model to be trained to obtain a plurality of target detection information output by the model to be trained;
[0019] Determining the recall rate and precision rate of the to-be-trained model corresponding to each of the initial parameter sets according to the plurality of preset feature matrices and the plurality of target detection information;
[0020] Based on the recall rate and the precision rate, multiple objective function values corresponding to the multiple initial parameter sets are obtained.
[0021] Optionally, determining the predicted mean and variance based on the multiple initial parameter sets and the objective function value includes:
[0022] Determining a target array formed by combining multiple initial parameter sets;
[0023] The predicted mean and the variance are determined based on the covariance vector between the next initial parameter set and all initial parameter sets in the target array, the covariance matrix between all initial parameter sets in the target array, and the target array.
[0024] Optionally, determining the next initial parameter set of the base classifier based on the predicted mean and variance includes:
[0025] Substitute the predicted mean and the variance into the expected improvement function to obtain the next initial parameter set of the base classifier.
[0026] Optionally, determining a target feature matrix of a target application according to the target feature information set includes:
[0027] Determining an initial feature matrix of the target application based on the target feature information set; wherein the initial feature matrix includes a tag value, and the tag value is used to distinguish malicious feature information and non-malicious feature information in the target feature information set;
[0028] Subtracting a target value from the tag value corresponding to the target application in the initial feature matrix to obtain a standard matrix; the target value is determined according to the tag value corresponding to the target application;
[0029] Determining a covariance matrix of the standard matrix according to the standard matrix and a transposed matrix of the standard matrix;
[0030] Solving the covariance matrix by a characteristic equation to obtain eigenvalues in the covariance matrix and eigenvectors corresponding to the eigenvalues;
[0031] Constructing a projection matrix according to the eigenvalues and the eigenvectors;
[0032] The target feature matrix is determined according to the projection matrix and the standard matrix.
[0033] Optionally, constructing a projection matrix according to the eigenvalues and the eigenvectors includes:
[0034] Sorting the characteristic values according to a preset rule to obtain a characteristic sequence;
[0035] Extracting a target feature value at a preset position from the feature sequence, and determining a target feature vector corresponding to the target feature value;
[0036] Use the target eigenvector as a column vector to construct an intermediate matrix;
[0037] A projection matrix is constructed according to the intermediate matrix and a transposed matrix of the intermediate matrix.
[0038] In a second aspect, a malicious application detection system is provided, the system comprising:
[0039] An acquisition module is configured to acquire a target feature information set of a target application; the target feature information set includes at least one of the following: sensitive permission features, system call features, and API call features of the target application;
[0040] a determination module, configured to determine a target feature matrix of a target application based on the target feature information set;
[0041] An output module is used to input the target feature matrix into a target model, so that a base classifier in the target model outputs a detection probability of the target application based on the target feature matrix, and a meta-classifier outputs detection information of the target application based on the detection probability and the target feature matrix; wherein the detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application; the base classifier includes at least two of the following: random forest RF, support vector machine SVM, logistic regression LR, and K-nearest neighbor algorithm KNN.
[0042] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, wherein the processor implements the detection method described in the first aspect when executing the computer program.
[0043] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the detection method described in the first aspect is implemented.
[0044] In a fifth aspect, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the detection method described in the first aspect is implemented.
[0045] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present application.
[0046] The above-mentioned malicious application detection method, system, device, medium, and product select at least two base classifiers from random forest RF, support vector machine SVM, logistic regression LR, and K-nearest neighbor algorithm KNN for the target model, ensuring that the selected base classifiers have significant differences and high efficiency and accuracy, reducing the impact of classification errors of a single classifier, while enhancing the target model's recognition ability for minority classes and improving the accuracy of judgment on imbalanced data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 1 is a flow chart of a method for detecting malicious applications in one embodiment;
[0048] Figure 2FIG1 is a schematic diagram of a dynamic analysis process of a method for detecting malicious applications in one embodiment;
[0049] Figure 3 Schematic diagram of a training process of a to-be-trained model of a method for detecting malicious applications in one embodiment;
[0050] Figure 4 1 is a schematic diagram of the structure of a malicious application detection system in one embodiment;
[0051] Figure 5 FIG. 1 is a schematic structural diagram of an electronic device in an embodiment. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0053] It should be noted that the diagrams provided in the present embodiment are only schematic illustrations of the basic concept of the present application. The diagrams only show the components related to the present application rather than the number, shape and size of the components when actually implemented. The type, quantity and ratio of each component can be changed at will during actual implementation, and the component layout pattern may also be more complicated. The structures, ratios, sizes, etc. illustrated in the drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read. They are not used to limit the restrictive conditions that can be implemented in this application. Therefore, they have no technical significance. Any modification of the structure, change of the proportional relationship or adjustment of the size should still fall within the scope of the technical content disclosed in this application without affecting the effect and purpose that can be achieved by this application. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of the implementation of this application. The change or adjustment of their relative relationship should also be considered as the scope of the implementation of this application without substantial change in the technical content.
[0054] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places herein does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0055] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include additional steps or elements.
[0056] The definition of inclusion herein, such as the terms “having”, “may have”, “include” or “may include” as used herein, indicates the existence of the corresponding functions, operations, elements, etc. herein, and does not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that the terms “including” or “having” as used herein indicate the existence of the features, numbers, steps, operations, elements, components or their combination described in the specification, and do not exclude the existence or addition of one or more other features, numbers, steps, operations, elements, components or their combination.
[0057] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a restriction on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary restriction. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "plurality" is two or more.
[0058] Figure 1 A method for detecting malicious applications provided in an exemplary embodiment of the present application includes:
[0059] S11. Obtain a target feature information set of a target application.
[0060] The target application may be an Android application, and the target application may be one or more.
[0061] The target application's feature information set includes multiple feature information extracted from the target application, and the target feature information set includes at least one of the following: sensitive permission features, system call features, and API (Application Programming Interface) call features of the target application. Sensitive permission features indicate permission features used to access user private information or perform certain sensitive operations, system call features indicate the characteristics and behaviors of interfaces (system calls) provided by the operating system kernel, and API call features indicate the characteristics and behaviors exhibited by the application programming interface during a call.
[0062] In order to obtain more comprehensive feature information, thereby comprehensively analyzing the characteristics and behaviors of the target application, improving the robustness of the detection method, improving the accuracy of malicious application detection, effectively identifying malicious applications, and preventing malicious applications from performing malicious behaviors, in one embodiment, obtaining a target feature information set of the target application includes:
[0063] Perform static analysis on the target application's installation package to obtain sensitive permission characteristics and API call characteristics;
[0064] Dynamically run the target application to obtain system call characteristics.
[0065] In the Android system, each application must declare the permissions it requires to perform operations. Malicious applications often request sensitive permissions that are irrelevant to their functionality in order to perform malicious actions. Therefore, by analyzing the permission declarations in the target application installation package, we can identify whether the application has requested unreasonable permissions. As shown in Table 1, Table 1 lists several sensitive permissions. These sensitive permissions may have a significant impact on the device running the target application and the privacy of the user using the device. Therefore, we can extract the sensitive permission characteristics in the target application for analysis.
[0066] Table 1 Sensitive permissions example table
[0067]
[0068]
[0069] In addition to sensitive permission features, API call features can also be used to analyze the behavior patterns of target applications and identify potential malicious behaviors.
[0070] The installation package of the target application can be statically analyzed to obtain the static features of the target application, such as sensitive permission features and API call features. Specifically, decompilation tools can be used, such as Apktool (APK decompilation tool) and dex2jar (Android dex to jar decompilation tool) to unpack the APK (Android Application Package, Android application package, also known as the installation package) of the target application, reversely convert the Android bytecode file into a Java bytecode file tool that can be statically analyzed and processed, decompile to obtain the source code file of the target application, preprocess the source code file, and then extract sensitive permission features and API call features from the preprocessed source code file.
[0071] AndroidManifest.xml (Android application manifest file) is a necessary file for every Android application. When the target application is packaged into an APK file, this file will be included in the APK. The file contains the basic metadata and configuration information of the target application. Therefore, in order to extract the system call features later, it is also necessary to obtain the source code file in the above static analysis stage.
[0072] AndroidManifest.xml, and extract some key information from AndroidManifest.xml, such as the package name of the target application (used to uniquely identify the target application) and the entry Activity information (the first interface displayed after the target application is started).
[0073] After completing static analysis of the target application and extracting sensitive permission features and API call features, we can further extract dynamic features of the target application, such as system call features. System calls are the interface through which the target application interacts with the operating system kernel. By making system calls, the target application can request the operating system to perform specific tasks, such as file operations, management, and communication. As shown in Table 2, which exemplifies multiple system calls, malicious applications often use system calls to perform malicious capabilities, such as stealing user data and destabilizing the system. Therefore, system call features can serve as important features for detecting malicious applications.
[0074] Table 2 System call example table
[0075] Serial number System calls 1 access 2 chmod 3 chown 4 open 5 ioctl 6 brkread 7 write 8 clone 9 close 10 exceve
[0076] In the dynamic analysis phase, Figure 2As shown, an automated test script can be generated based on the package name and entry activity information extracted from the static analysis above. Running this automated test script can simulate various user interaction scenarios with the target application, such as clicking buttons, entering text, and swiping the screen. The target application's APK is then installed on the emulator using the adb install command (a command-line tool used to install applications on Android devices). After successful installation, the target application enters the automated testing phase. During this phase, the MonkeyRunner tool (a tool for automated testing of Android applications) can be used to run pre-written automated test scripts. This simulates various user interaction scenarios with the target application, maximizing coverage of all the target application's functionalities and behavior paths. Using the MonkeyRunner tool to test scripts, a large number of simulated events can be automatically executed. MonkeyRunner's timer can also be set. When the automated test script reaches its execution time or all test items in the automated test script have been analyzed, system call signatures are output. The target signature information set is composed of system call signatures, API call signatures, and sensitive permission signatures.
[0077] By performing dynamic and static analysis on the target application's installation package, we can effectively analyze the target application's code structure, visually display its actual behavior, and obtain more comprehensive feature information. By fusing these features, we can more comprehensively analyze the characteristics and behavior of the target application. This improves the accuracy of malicious application detection, effectively identifies malicious applications, and prevents them from executing malicious behaviors.
[0078] S12. Determine a target feature matrix of the target application based on the target feature information set.
[0079] In one embodiment, determining a target feature matrix of a target application based on the target feature information set may include:
[0080] After extracting a target feature information set including multiple sensitive permission features, multiple API call features, and multiple system call features from the target application, the various features in the target feature information set can be linearly fused first, and then the various features after linear fusion can be sorted by category. The features in the same category are arranged in lexicographic order, thereby forming a sorted target feature information set. The malicious feature information in the sorted target feature information set is then marked with a first tag value, for example, the first tag value is 1, and the non-malicious features are marked with a second tag value, for example, 0, for example, the second tag value is 0, thereby representing the target feature information set as a target feature matrix composed of 1 and 0. Malicious feature information is feature information that may be called by malicious applications and perform malicious behaviors. It is understandable that benign applications may also call malicious feature information but will not perform malicious behaviors.
[0081] In order to reduce the computational complexity of the target model and improve the computational efficiency of the target model, in another embodiment, determining the target feature matrix of the target application based on the target feature information set may include:
[0082] Determining an initial feature matrix of the target application based on the target feature information set; wherein the initial feature matrix includes a tag value, and the tag value is used to distinguish malicious feature information from non-malicious feature information in the target feature information set;
[0083] The target value is subtracted from the tag value corresponding to the target application in the initial feature matrix to obtain a standard matrix; the target value is determined according to the tag value corresponding to the target application;
[0084] Determine the covariance matrix of the standard matrix based on the standard matrix and the transposed matrix of the standard matrix;
[0085] Solving the covariance matrix by a characteristic equation to obtain eigenvalues in the covariance matrix and eigenvectors corresponding to the eigenvalues;
[0086] Construct the projection matrix based on the eigenvalues and eigenvectors;
[0087] Determine the target feature matrix based on the projection matrix and the standard matrix.
[0088] After extracting a target feature information set including multiple sensitive permission features, multiple API call features, and multiple system call features from the target application, the individual features in the target feature information set can be linearly fused, and then the linearly fused features are sorted by category. Features in the same category are arranged in lexicographic order, thereby forming a sorted target feature information set. Malicious feature information in the sorted target feature information set is then marked with a first tag value, such as 1, and non-malicious features are marked with a second tag value, such as 0, such as 0, thereby representing the target feature information set as an initial feature matrix consisting of 1s and 0s.
[0089] Then calculate the mean for each row in the initial feature matrix. When the target application is one, the initial feature matrix is a row, and the first label value and / or the second label value in the row are subtracted from the mean of the row to obtain the standard matrix after de-averaging. When the target application is multiple applications, the initial feature matrix is multiple rows, and the first label value and / or the second label value in each row are subtracted from the mean of the row to obtain the standard matrix after de-averaging. Then according to the formula K=MM τ , calculate the covariance matrix K of the standard matrix M, where M is the standard matrix, M τ is the transposed matrix of the standard matrix, and K is the covariance matrix. Perform eigenvalue decomposition on the covariance matrix K using the characteristic equation K-λI=0 to obtain its eigenvalues λ and corresponding eigenvectors, where λ is the eigenvalue and I is the identity matrix. A projection matrix is then constructed based on the eigenvalues and eigenvectors. Finally, the product of the projection matrix and the standard matrix is calculated to obtain the target characteristic matrix.
[0090] The importance of an eigenvector is reflected by its corresponding eigenvalue. The larger the eigenvalue, the greater the variance of the data on the eigenvector, that is, the more significant the data change in that direction. Therefore, in order to obtain more important eigenvectors and construct a projection matrix based on the important eigenvectors, in one embodiment, the projection matrix is constructed based on the eigenvalues and eigenvectors, including:
[0091] Sort the eigenvalues according to preset rules to obtain a feature sequence;
[0092] Extracting a target feature value at a preset position from the feature sequence, and determining a target feature vector corresponding to the target feature value;
[0093] Use the target eigenvector as a column vector to construct an intermediate matrix;
[0094] Construct the projection matrix based on the intermediate matrix and the transposed matrix of the intermediate matrix.
[0095] First, the eigenvalues obtained after solving the above covariance matrix K are arranged in order from large to small, and then the first k eigenvalues are extracted as the target eigenvalues, and the target eigenvector corresponding to the target eigenvalue is determined. The target eigenvector is used as the column vector to construct the intermediate matrix. Finally, the product of the intermediate matrix and the transposed matrix of the intermediate matrix is calculated to obtain the projection matrix.
[0096] Principal component analysis is used to average and reduce the dimensionality of the initial feature matrix. Eigenvectors and eigenvalues are ranked according to their importance, retaining the effective components to achieve data dimensionality reduction and obtain the target feature matrix. This reduces the dimensionality of the target feature matrix, reducing computational complexity and improving the computational efficiency of the target model while retaining the most important information and improving the accuracy of the target model.
[0097] S13. Input the target feature matrix into the target model, so that the base classifier in the target model outputs the detection probability of the target application according to the target feature matrix, and the meta-classifier outputs the detection information of the target application according to the detection probability and the target feature matrix.
[0098] The detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application. The detection information may indicate that the target application is a malicious application or a non-malicious application.
[0099] The base classifiers include at least two of the following: RF (Random Forest), SVM (Support Vector Machine), LR (Logistic Regression), and KNN (K-Nearest Neighbors). Two or three of these base classifiers can be selected and added to the target model. More preferably, the four base classifiers are selected and added to the target model.
[0100] The selected base classifiers all have high detection accuracy and are relatively efficient. Higher accuracy and efficiency of a single base classifier generally significantly improve the overall performance and detection accuracy of the target model. Furthermore, the working principles of each base classifier vary significantly. While maintaining an acceptable error rate for the base classifiers, greater diversity between classifiers further enhances the accuracy of the target model's output. Selecting at least two base classifiers reduces the impact of individual classifier errors, enhances the target model's ability to recognize minority classes, and improves the accuracy of judgments for imbalanced datasets.
[0101] The meta-classifier may be XGBoost (eXtreme Gradient Boosting).
[0102] For imbalanced data sets, at least two base classifiers are selected, and the detection probability output by the base classifier and the target feature matrix are used as the input of the meta-classifier to reduce the impact of classification errors of a single classifier, while enhancing the target model's ability to recognize minority classes and improving the judgment accuracy of the imbalanced data set.
[0103] In order to optimize the parameters of the base classifiers in the target model, an optimal parameter set consisting of parameter values of multiple base classifiers is found. In one embodiment, the parameter set of the base classifiers is obtained by the following steps:
[0104] Determine an initial parameter set of a base classifier; the initial parameter set includes a plurality of initial parameter values;
[0105] Substituting multiple initial parameter sets into the model to be trained to obtain multiple objective function values corresponding to the multiple initial parameter sets;
[0106] Determine the prediction mean and variance based on multiple initial parameter sets, initial parameter values, and objective function values;
[0107] Based on the predicted mean and variance, the next initial parameter set of the base classifier is determined until the preset iteration stopping condition is met to obtain the parameter set of the base classifier.
[0108] In one embodiment, multiple initial parameter sets are substituted into the model to be trained to obtain multiple objective function values corresponding to the multiple initial parameter sets, including:
[0109] Substitute multiple initial parameter sets into the model to be trained respectively;
[0110] Corresponding to each initial parameter set, multiple preset feature matrices are input into the model to be trained to obtain multiple target detection information output by the model to be trained;
[0111] Determine the recall rate and precision rate of the to-be-trained model corresponding to each initial parameter set according to multiple preset feature matrices and multiple target detection information;
[0112] Based on the recall rate and precision rate, multiple objective function values corresponding to multiple initial parameter sets are obtained.
[0113] In one embodiment, determining a prediction mean and variance based on multiple initial parameter sets, initial parameter values, and objective function values includes:
[0114] Determine a target array composed of multiple initial parameter sets;
[0115] The predicted mean and variance are determined based on the covariance vector between the next initial parameter set and all initial parameter sets in the target array, the covariance matrix between all initial parameter sets in the target array, the target array, and the objective function value.
[0116] like Figure 3 As shown, in order to obtain the parameter set of the base classifier, the test set and sample set can be obtained for the model to be trained. The steps for obtaining the test set and sample set are as follows: First, the installation packages of multiple applications can be statically analyzed to obtain static features of multiple applications, such as sensitive permission features and API call features. Specifically, decompilation tools can be used, such as Apktool and dex2jar to unpack the APKs of multiple applications, reverse convert the Android bytecode files into Java bytecode files that can be statically analyzed and processed, decompile to obtain the source code files of multiple applications, preprocess the source code files, and then extract sensitive permission features and API call features from the preprocessed source code files.
[0117] AndroidManifest.xml is a required file for every Android application. When an application is packaged into an APK file, this file is included in the APK and contains basic metadata and configuration information for the application. Therefore, in order to subsequently extract system call signatures, the aforementioned static analysis phase also requires obtaining AndroidManifest.xml from the source code files and extracting key information from it, such as the application's package name (which uniquely identifies the application) and entry point Activity information (the first screen displayed after the application is launched).
[0118] After completing static analysis of multiple applications and extracting sensitive permission signatures and API call signatures, further dynamic signatures, such as system call signatures, can be extracted. An automated test script can be generated based on the package name and entry activity information extracted from the static analysis. Running this automated test script simulates various user interaction scenarios with multiple applications, such as clicking buttons, entering text, and swiping the screen. The APKs of the multiple applications are then installed on the simulator using the adb install command. After successful installation, the multiple applications enter the automated testing phase. During this phase, pre-written automated test scripts can be run using the MonkeyRunner tool. These scripts simulate various user interaction scenarios, maximizing coverage of all functional and behavioral paths across multiple applications. Using MonkeyRunner to test scripts, a large number of simulated events can be automatically executed. MonkeyRunner timers can also be set. When the automated test script reaches its execution time or when all test items in the automated test script have been analyzed, system call signatures are output.
[0119] According to the above steps, a feature information set consisting of α sensitive permission features, γ system call features and β API call features of multiple applications is obtained. The feature vectors of sensitive permission features, API call features and system call features are [p1, p2, ..., pα], [a1, a2, ..., aβ], [i1, i2, ..., iγ] respectively. The sensitive permission feature vector is denoted as [P], the API call feature vector is denoted as [A], and the system call feature vector is denoted as [I]. Then, each feature in the feature information set is preprocessed, and the applications that cannot extract features or features that appear too few times in the feature information set are deleted to reduce their impact on the detection results. Then, each feature in the feature information set is linearly fused, and then the linearly fused features are sorted according to category. The features in the same category are arranged in lexicographic order to form a sorted feature information set. The feature information set C consists of the sensitive permission feature vector [P], the API feature vector [A] and the system call feature vector [I]. Its dimension is α+β+γ. The feature information set C = [p1, p2, ..., p α , a1, a2, ..., a β ,i1,i2,…,i γ ]. Then the malicious feature information in the sorted feature information set is marked with a third mark value, for example, the third mark value is 1, and the non-malicious feature is marked with a fourth mark value, for example, 0, for example, the fourth mark value is 0, thereby representing the feature information set as a first feature matrix composed of 1 and 0.
[0120] Then calculate the mean of each row in the first feature matrix, subtract the mean from the third label value and / or the fourth label value in each row to obtain the standard matrix after de-averaging. Then according to the formula K=MM τ , calculate the covariance matrix K of the standard matrix M, where M is the standard matrix, M τ is the transposed matrix of the standard matrix, and K is the covariance matrix. Perform eigenvalue decomposition on the covariance matrix K to obtain its eigenvalues and corresponding eigenvectors. The eigenvalues obtained from solving the covariance matrix K are then arranged in descending order. The eigenvectors corresponding to the first n eigenvalues are then used as column vectors to construct an intermediate matrix. Finally, the product of the intermediate matrix and its transposed matrix is calculated to obtain the projection matrix. Finally, the product of the projection matrix and the standard matrix is calculated to obtain the second characteristic matrix.
[0121] After obtaining the second feature matrix, the second feature matrix can be divided into a test set and a sample set according to a certain ratio. For example, the second feature matrix can be divided into a sample set and a test set according to a ratio of 7:3. The sample set is input into the model to be trained, and the test set is used to verify the generalization ability and accuracy of the model to be trained.
[0122] After obtaining the test set and sample set, the following describes how to obtain the parameter set of the base classifier:
[0123] The model to be trained in this application can be a Gaussian process regression model, and the target model can be a Stacking ensemble learning model. The base classifiers in the Stacking ensemble learning model can include at least two of RF, SVM, LR, and KNN. The output value and sample set of the base classifier are then input into the meta-classifier XGBoost, which outputs the detection information of multiple applications based on the output value and sample set of the base classifier. The differences between the four base classifiers are large and relatively efficient, which can improve the accuracy of the output of the Stacking ensemble learning model.
[0124] As shown in Table 3, the kernel function Kernel and gamma value in the base classifier SVM can help the trained model find a suitable decision boundary, and the C value controls the complexity of the trained model. The number of neighbors n_neighbors in KNN determines the sensitivity of classification. The number of trees n_estimators and the maximum depth max_depth in RF balance the stability and complexity of the trained model. The regularization type penalty in LR affects feature selection and the generalization ability of the trained model.
[0125] Table 3 Key parameters of base classifier
[0126]
[0127] Therefore, these key parameters can be combined and tuned. Each parameter takes a value to form different parameter sets. First, a set of parameter combinations is input into the model to be trained. As the sampling points gradually increase, the prediction results of the model to be trained will become closer and closer to the actual results.
[0128] As shown in Table 4, we can first select a set of initial parameter sets xn of the base classifier by random method:
[0129] Table 4 Initial parameter set
[0130]
[0131] Then, according to the above method, multiple initial parameter sets are selected, with x1 to xn being each initial parameter set, and a target array X0 = {x1, x2, ..., xn} formed by combining the multiple initial parameter sets is input into the model to be trained. Multiple preset feature matrices, i.e., sample sets, are input into the model to be trained to obtain multiple target detection information output by the model to be trained. Based on whether the target detection information is consistent with the benign and malignant applications in the sample set, the recall rate and precision rate corresponding to each initial parameter set of the model to be trained are calculated. According to the calculation formulas for recall rate, precision rate, and F1 value, the objective function value = 2 * (recall rate * precision rate) / (recall rate + precision rate) is calculated, and multiple objective function values y1 to yn corresponding to the multiple initial parameter sets are obtained. The multiple objective function values are combined to form Y0 = {y1, y2, ..., yi}, where yi = f(xi) and f(x) is the objective function.
[0132] Next, the observed data points (X0, Y0) are used to update the posterior distribution of the Gaussian process and train the model to be trained. Assuming that the objective function f(x) obeys the Gaussian process, the initial parameter set x of the base classifier is set, the mean function of the objective function is m(x), and the covariance function is k(x, x′), then the predicted mean μ(x) and variance σ are 2 (x) are:
[0133] μ(x)=m(x)+k(x,X0)K(X0,X0) -1 (Y0-m(X0));
[0134] σ 2 (x)=k(x,x)-k(x,X0)K(X0,X0) -1 k(X0, x);
[0135] Among them, k(x, X0) is the covariance vector between the next initial parameter set x of the base classifier and all initial parameter sets in the target array X0, and K(X0, X0) is the covariance matrix between all initial parameter sets in the target array X0.
[0136] Get the predicted mean μ(x) and variance σ 2 After (x), in one embodiment, the next initial parameter set of the base classifier is determined based on the predicted mean and variance, including:
[0137] Substitute the predicted mean and variance into the expected improvement function to obtain the initial parameter set of the base classifier for the next time.
[0138] The acquisition function a(x) can select the next initial parameter set that is most likely to improve the current optimal objective function value, i.e., the F1 value, based on the posterior distribution. Considering the imbalanced dataset in this application, this application selects the expected improvement function as the acquisition function.
[0139] Because EI (Expected Improvement) can find a good balance between exploration (exploration refers to searching for new, unevaluated areas in the parameter space to discover parameter combinations that may have higher F1 values) and exploitation (exploitation refers to a more detailed search near the initial parameter set with known good performance to further improve the F1 value), it considers both the probability of improvement (the probability of improvement refers to the probability that the next parameter set will bring about performance improvement based on the current optimal solution, and EI is measured by the predicted mean and uncertainty) and the expected size of the improvement (the expected size of the improvement refers to the expected value of the performance improvement that the next parameter combination can bring about based on the current optimal solution, and EI is reflected by calculating the expected improvement value). Because the F1 value is usually an indicator that comprehensively considers precision and recall, it is necessary to find a parameter set that can improve both aspects simultaneously, so this is particularly important for the F1 value optimization problem.
[0140] Therefore, the EI function can be used. The formula of the EI function is as follows:
[0141]
[0142] Among them, μ(x) and σ(x) are the predicted mean and standard deviation of the model to be trained at x, respectively, and f * is the current optimal F1 value, Φ and σ are the cumulative distribution function and probability density function of the standard normal distribution, respectively.
[0143] According to the above formula, select the parameter combination x that maximizes the acquisition function a(x) next And determine it as the initial parameter set of the base classifier next time, run the model to be trained to get the new objective function value y next . The new observation data (x next ,y next) is added to the observation dataset (X0, Y0) to obtain a new observation dataset (X1, Y1). This process continues until a preset iteration stop condition is met, for example, when a preset number of iterations is reached, 50, or a parameter combination that meets the requirements is found, the iteration is stopped, and the trained target model and the parameter set of the base classifier of the target model are obtained.
[0144] After obtaining the trained target model, it can receive new sample set input and use the trained target model for classification. After sufficient training and optimization, the target model can accurately identify malicious and benign Android applications, helping users defend against potential security threats in real time.
[0145] This application uses the Bayesian optimization method to optimize the parameter set of the base classifier of the target model, and selects F1 as the target value for the imbalanced data set, which greatly improves the generalization ability of the model.
[0146] It should be understood that although Figure 1-3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1-3 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0147] like Figure 4 As shown, the present application also provides a malicious application detection system, the detection system comprising:
[0148] An acquisition module 41 is configured to acquire a target feature information set of a target application; the target feature information set includes at least one of the following: sensitive permission features, system call features, and API call features of the target application;
[0149] A determination module 42 is configured to determine a target feature matrix of a target application based on the target feature information set;
[0150] An output module 43 is configured to input the target feature matrix into a target model, so that a base classifier in the target model outputs a detection probability of the target application based on the target feature matrix, and a meta-classifier outputs detection information of the target application based on the detection probability and the target feature matrix; wherein the detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application; and the base classifier includes at least two of the following: random forest RF, support vector machine SVM, logistic regression LR, and K-nearest neighbor algorithm KNN.
[0151] In one embodiment, the acquisition module 41 is further configured to:
[0152] Performing static analysis on the installation package of the target application to obtain the sensitive permission characteristics and the API call characteristics;
[0153] The target application is dynamically run to obtain the system call characteristics.
[0154] In one embodiment, the parameter set of the base classifier is obtained by the following steps:
[0155] Determining an initial parameter set of the base classifier;
[0156] Substituting the multiple initial parameter sets into the to-be-trained model to obtain multiple objective function values corresponding to the multiple initial parameter sets;
[0157] Determining a predicted mean and variance based on the multiple initial parameter sets and the objective function value;
[0158] Based on the predicted mean and variance, the next initial parameter set of the base classifier is determined until a preset iteration stop condition is met, thereby obtaining the parameter set of the base classifier of the target model.
[0159] In one embodiment, the multiple initial parameter sets are substituted into the model to be trained to obtain multiple objective function values corresponding to the multiple initial parameter sets, including:
[0160] Substituting the multiple initial parameter sets into the to-be-trained model respectively;
[0161] Corresponding to each initial parameter set, inputting a preset feature matrix into the model to be trained to obtain a plurality of target detection information output by the model to be trained;
[0162] Determining the recall rate and precision rate of the to-be-trained model corresponding to each of the initial parameter sets according to the plurality of preset feature matrices and the plurality of target detection information;
[0163] Based on the recall rate and the precision rate, multiple objective function values corresponding to the multiple initial parameter sets are obtained.
[0164] In one embodiment, determining the predicted mean and variance based on the multiple initial parameter sets and the objective function value includes:
[0165] Determining a target array formed by combining multiple initial parameter sets;
[0166] The predicted mean and the variance are determined based on the covariance vector between the next initial parameter set and all initial parameter sets in the target array, the covariance matrix between all initial parameter sets in the target array, and the target array.
[0167] In one embodiment, determining the next initial parameter set of the base classifier based on the predicted mean and variance includes:
[0168] Substitute the predicted mean and the variance into the expected improvement function to obtain the next initial parameter set of the base classifier.
[0169] In one embodiment, the determination module 42 is further configured to:
[0170] Determining an initial feature matrix of the target application based on the target feature information set; wherein the initial feature matrix includes a tag value, and the tag value is used to distinguish malicious feature information and non-malicious feature information in the target feature information set;
[0171] Subtracting a target value from the tag value corresponding to the target application in the initial feature matrix to obtain a standard matrix; the target value is determined according to the tag value corresponding to the target application;
[0172] Determining a covariance matrix of the standard matrix according to the standard matrix and a transposed matrix of the standard matrix;
[0173] Solving the covariance matrix by a characteristic equation to obtain eigenvalues in the covariance matrix and eigenvectors corresponding to the eigenvalues;
[0174] Constructing a projection matrix according to the eigenvalues and the eigenvectors;
[0175] The target feature matrix is determined according to the projection matrix and the standard matrix.
[0176] In one embodiment, the determination module 42 is further configured to:
[0177] Sorting the characteristic values according to a preset rule to obtain a characteristic sequence;
[0178] Extracting a target feature value at a preset position from the feature sequence, and determining a target feature vector corresponding to the target feature value;
[0179] Use the target eigenvector as a column vector to construct an intermediate matrix;
[0180] A projection matrix is constructed according to the intermediate matrix and a transposed matrix of the intermediate matrix.
[0181] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only illustrative, in which the units described as separate components may or may not be physically separated, and the components of the units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application solution.
[0182] Figure 5 This is a structural diagram of an electronic device shown in an example embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor. When the processor executes the computer program, it implements the malicious application detection method described in any of the above embodiments. Figure 5 The electronic device 50 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0183] like Figure 5 As shown, the electronic device 50 may be a general-purpose computing device, such as a server device. Components of the electronic device 50 may include, but are not limited to, the at least one processor 51, the at least one memory 52, and a bus 53 connecting different system components (including the memory 52 and the processor 51).
[0184] The bus 53 includes a data bus, an address bus, and a control bus.
[0185] The memory 52 may include a volatile memory, such as a random access memory (RAM) 521 and / or a cache memory 522 , and may further include a read-only memory (ROM) 523 .
[0186] The memory 52 may also include a program tool 525 (or utility) having a set (at least one) of program modules 524, such program modules 524 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0187] The processor 51 executes various functional applications and data processing by running computer programs stored in the memory 52, such as the malicious application detection method provided in any of the above embodiments.
[0188] The electronic device 50 can also communicate with one or more external devices 54 (e.g., a keyboard, pointing device, etc.). Such communication can be performed via an input / output (I / O) interface 55. Furthermore, the electronic device 50 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 56. As shown, the network adapter 56 communicates with other modules of the electronic device 50 via a bus 53. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 50, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0189] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0190] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the malicious application detection method provided in any of the above embodiments.
[0191] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0192] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0193] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for detecting malicious applications.
[0194] The program code for executing the computer program product of the present application may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0195] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0196] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for detecting malicious applications, characterized in that: The detection method comprises: Obtain a target feature information set of a target application; the target feature information set includes at least one of the following: sensitive permission features, system call features, and API call features of the target application; Determining a target feature matrix of a target application according to the target feature information set; The target feature matrix is input into the target model, so that the base classifier in the target model outputs the detection probability of the target application according to the target feature matrix, and the meta-classifier outputs the detection information of the target application according to the detection probability and the target feature matrix; wherein the detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application; the base classifier includes at least two of the following: random forest RF, support vector machine SVM, logistic regression LR, K-nearest neighbor algorithm KNN.
2. The detection method according to claim 1, characterized in that The step of obtaining a target feature information set of a target application includes: Performing static analysis on the installation package of the target application to obtain the sensitive permission characteristics and the API call characteristics; The target application is dynamically run to obtain the system call feature.
3. The detection method according to claim 1, characterized in that The parameter set of the base classifier is obtained by the following steps: Determining an initial parameter set of the base classifier; Substituting the multiple initial parameter sets into the model to be trained to obtain multiple objective function values corresponding to the multiple initial parameter sets; Determining a predicted mean and variance according to the multiple initial parameter sets and the objective function value; Based on the predicted mean and variance, the next initial parameter set of the base classifier is determined until a preset iteration stop condition is met, thereby obtaining a parameter set of the base classifier of the target model.
4. The detection method according to claim 3, characterized in that Substituting the multiple initial parameter sets into the model to be trained to obtain multiple objective function values corresponding to the multiple initial parameter sets includes: Substituting the multiple initial parameter sets into the models to be trained respectively; Corresponding to each initial parameter set, inputting a preset feature matrix into the model to be trained to obtain a plurality of target detection information output by the model to be trained; According to the plurality of preset feature matrices and the plurality of target detection information, respectively determining the recall rate and the precision rate of the to-be-trained model corresponding to each of the initial parameter sets; Based on the recall rate and the precision rate, multiple objective function values corresponding to the multiple initial parameter sets are obtained.
5. The detection method according to claim 3, characterized in that: Determining the predicted mean and variance according to the multiple initial parameter sets and the objective function value includes: Determine a target array formed by combining a plurality of the initial parameter sets; The predicted mean and the variance are determined based on the covariance vector between the next initial parameter set and all initial parameter sets in the target array, the covariance matrix between all initial parameter sets in the target array, and the target array.
6. The detection method according to claim 3, characterized in that: The step of determining the next initial parameter set of the base classifier based on the predicted mean and variance includes: Substitute the predicted mean and the variance into the expected improvement function to obtain the next initial parameter set of the base classifier.
7. The detection method according to claim 1, characterized in that The step of determining a target feature matrix of a target application according to the target feature information set includes: Determine an initial feature matrix of the target application according to the target feature information set; wherein the initial feature matrix includes a tag value, and the tag value is used to distinguish malicious feature information and non-malicious feature information in the target feature information set; Subtracting the target value from the tag value corresponding to the target application in the initial feature matrix to obtain a standard matrix; the target value is determined according to the tag value corresponding to the target application; Determine a covariance matrix of the standard matrix according to the standard matrix and a transposed matrix of the standard matrix; Solving the covariance matrix by means of a characteristic equation to obtain eigenvalues in the covariance matrix and eigenvectors corresponding to the eigenvalues; Constructing a projection matrix according to the eigenvalues and the eigenvectors; The target feature matrix is determined according to the projection matrix and the standard matrix.
8. The detection method according to claim 7, characterized in that The step of constructing a projection matrix according to the eigenvalue and the eigenvector comprises: Sorting the characteristic values according to a preset rule to obtain a characteristic sequence; Extracting a target feature value at a preset position from the feature sequence, and determining a target feature vector corresponding to the target feature value; Use the target feature vector as a column vector to construct an intermediate matrix; A projection matrix is constructed according to the intermediate matrix and a transposed matrix of the intermediate matrix.
9. A malicious application detection system, characterized in that: The detection system comprises: An acquisition module, used to acquire a target feature information set of a target application; the target feature information set includes at least one of the following: sensitive permission features, system call features, and API call features of the target application; A determination module, used to determine a target feature matrix of a target application according to the target feature information set; An output module is used to input the target feature matrix into a target model, so that a base classifier in the target model outputs a detection probability of the target application according to the target feature matrix, and a meta-classifier outputs detection information of the target application according to the detection probability and the target feature matrix; wherein the detection probability indicates the probability that the target application is a malicious application, and the detection information indicates whether the target application is a malicious application; the base classifier includes at least two of the following: random forest RF, support vector machine SVM, logistic regression LR, K-nearest neighbor algorithm KNN.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the detection method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the detection method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the detection method according to any one of claims 1 to 8 is implemented.