APP fraud-related detection method based on API semantic analysis
Through detection methods based on API semantic analysis, combined with feature refinement, semantic generation and word embedding fusion, the risk problem in existing API detection methods is solved in the difficulty of identifying API combination and semantic transformation, and more accurate API behavior analysis and malicious detection are achieved.
Patent Information
- Application Number
- CN202510195752.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
Existing API detection methods are difficult to accurately identify potential risks in API combinations and semantic transformations, lack of understanding of the context, and are difficult to deal with complex malicious behavior.
Using detection methods based on API semantic analysis, API-specific semantic word embedding is constructed through feature refinement, semantic generation and word embedding fusion, and fusion with the BERT model for fine-tuning to generate API feature statements containing rich context.
It improves the understanding and analysis ability of API use, enhances the ability to capture complex behavior patterns, and improves the generalization ability of the model and the precise understanding of API behavior.
Smart Images

Figure CN120030541A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of software security detection and relates to an APP fraud detection method based on API semantic analysis. Background Art
[0002] With the rapid development of mobile Internet, smartphones and various mobile applications (APPs) have become an important part of people's daily life and work. From basic social, shopping, and payment to complex financial, medical, and enterprise management application scenarios, the functions of APPs are constantly expanding, and they are deeply bound to users' privacy data, system permissions, and network communications. However, with the prosperity of the APP ecosystem, security risks are becoming increasingly prominent, especially in terms of permission abuse, malicious API calls, and privacy data leakage. Attackers often use APIs to perform unauthorized operations, dynamically load malicious code, or evade detection mechanisms, which poses a huge challenge to security protection. In the current field of mobile security research, how to accurately analyze API semantics and explore the correlation between APIs to effectively detect malicious behavior has become one of the core issues that need to be solved urgently.
[0003] However, traditional API detection methods mostly rely on directly inputting API names and parameter information in terms of feature extraction. The features generated based on this are relatively simple. In malicious behavior detection tasks, attackers often hide malicious intentions through API combinations or semantic transformations. Traditional methods find it difficult to accurately identify potential risks.
[0004] Traditional methods mainly rely on static analysis and dynamic analysis combined with rule matching and behavior monitoring to identify the characteristics of malicious applications. Traditional detection methods such as static detection based on feature rules, dynamic detection based on behavior analysis, and detection methods based on machine learning have limitations such as difficulty in dealing with API combinations and semantic changes, and lack of understanding of context. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide an APP fraud detection method based on API semantic analysis. It is an API detection method that combines feature refinement, semantic generation and word embedding fusion. Feature refinement mainly refines features from several aspects, including the relationship between API parameters and APIs, the relationship between APIs, API call frequency, API call sequence and combination mode, and functional description of APIs and parameters. Semantic generation uses features to generate API semantic statements, and constructs an API-specific semantic corpus through API documents, generates API-specific semantic word embeddings, and fuses them with BERT and training word embeddings. Finally, the API semantic statements are used to fine-tune the BERT model.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for detecting APP fraud based on API semantic analysis, the method comprising the following steps:
[0008] S1. Obtain APP samples, and perform static and dynamic analysis on the APP samples to obtain parameter information and operation logs of the APP samples, and perform refined feature analysis on the parameter information of the APP samples to construct API feature statements, and build an API corpus based on the APP sample operation logs;
[0009] S2. Based on the API corpus and API feature sentences, use Word2Vec training to generate API-specific semantic word vectors;
[0010] S3. Combine the context vector generated by BERT for weighted fusion to obtain the fused word vector, and embed the fused word vector into the word embedding layer of the BERT pre-trained model;
[0011] S4. Fine-tune the BERT model in the word embedding layer. Use API feature statements to fine-tune the weights of the Embedding layer and Transformer layer of the BERT model. After the training is completed, predict the fraud-related malicious apps.
[0012] Further, in step S1, it includes:
[0013] S11. Obtain APP samples, obtain preliminary features through dynamic and static analysis, and perform detailed analysis on API features to finally obtain API feature statements;
[0014] S12. Preliminarily build an API corpus by collecting the application's API call logs and API documents.
[0015] Further, in step S11, by analyzing the API and its API parameters, and the APP behavior monitoring log, multi-dimensional features are obtained, and the multi-dimensional features include: the association relationship between APIs, API function description, API call frequency, and API call sequence; and then the multi-dimensional features are used to generate rules for template-filled feature generation to generate API feature statements with natural language logic and natural language descriptions of refined features; refined feature generation includes:
[0016] By analyzing the parameters of an API, the association relationship between the API and other APIs is obtained, and a feature statement of the association relationship between the API and other APIs is generated; wherein, the input and output parameters of a certain API are first analyzed, and the association relationship between its parameters and other APIs is extracted, and a feature statement is generated according to the rules of template filling feature generation through the association relationship between APIs;
[0017] Obtaining API function descriptions and generating feature statements related to the API function descriptions; wherein, obtaining the core function descriptions of the API from the description field in the API document and directly generating feature statements based on the function descriptions;
[0018] Analyze the calling frequency of APIs and generate feature statements about the calling frequency of APIs; by analyzing the application running log or API calling stack, count the number of calls and calling frequency of each API in different scenarios, mark APIs with abnormally high or low frequency calls, and generate feature statements according to the calling frequency according to the rules of template filling feature generation;
[0019] Analyze the calling sequence of APIs and generate characteristic statements about API calls. Based on the calling relationship graph of the application, record the dependencies and calling paths of API calls to form calling sequence information that reflects the logical process, so as to capture the logical characteristics of malicious behavior and construct a malicious calling relationship graph, in which each node represents an API call and each edge represents the calling sequence relationship. Extract the API calling path and calculate the calling frequency. For the calling path with high calling frequency, encode the calling path into a serialized text description and generate characteristic statements according to the calling relationship.
[0020] Further, in step S12, the process of constructing the API corpus is:
[0021] First, extract the key fields of the API call from the call log, including the API name, parameter information, and return value;
[0022] Then, parse the API document content to extract the API function description, parameter description, and applicable scenario information;
[0023] Then clean the text information, remove special symbols and unify the capitalization;
[0024] Finally, the processed text is segmented to obtain the corpus ApiLex containing API terms.
[0025] Further, in step S2, the process of using Word2Vec training to generate API-specific semantic word vectors is:
[0026] Use the unsupervised training method Skip-gram algorithm to train the Word2Vec model;
[0027] Adjust the settings before training. Set the window size that determines the range of the "window" when the model captures the context of the word, set the word vector dimension to determine the dimension of the vector representation of each word, and set the minimum word frequency threshold to ignore words with a frequency lower than the minimum word frequency threshold.
[0028] During the training process, for a text sample, the model first initializes vectors for the words in the text sample and assigns a randomly initialized vector to each word. Then, in each training iteration, the context sliding window slides on the text sample. During the sliding process, each word in the text sample is used as the target word in the window, and the upper and lower words of the target word are generated according to the window size. The loss function is calculated based on the gap between the generated upper and lower words and the actual upper and lower words. After multiple iterations, when the loss function converges, the training is completed.
[0029] After training is completed, API-specific semantic word vectors are obtained based on the API corpus and API feature sentences.
[0030] Furthermore, in step S3, the API-specific word vectors generated by Word2Vec and the word vectors pre-trained by BERT are aligned, and the high-dimensional vectors are reduced in dimension using dimensionality reduction technology. Then, a weighted fusion strategy is used to assign weights to the two types of vectors. The weights are optimized according to specific tasks to obtain the fused word vectors:
[0031] V f =αV w +βV b
[0032] Where V f Represents the fused word vector, V w Represents the word vector generated by the Word2Vec model, V b Represents the word vector pre-trained by the BERT model.
[0033] Furthermore, embedding the fused word vector into the word embedding layer of the BERT pre-trained model means: setting the fused word vector as each input subword token in the Embedding layer of the BERT pre-trained model.
[0034] Further, in step S4, fine-tuning training for the BERT model includes the following steps:
[0035] S41. Improve the data set. The APP feature statements of the sample APPs will be used as model inputs. The APP samples will be labeled according to whether the APPs are involved in fraud, and a complete training data set will be obtained.
[0036] S42. Optimize the data set, cluster multiple feature sentences of each sample, obtain multiple representative feature sentences, and use the DBSCAN clustering method to cluster the API sentence embeddings;
[0037] First, obtain the API statement of the existing weighted fusion model, E = {E 1 ,E 2 ,…,E N}, select cosine similarity as the similarity measure in the clustering process, configure DBSCAN hyperparameters, embed the sentence E as the input of DBSCAN, perform density calculation and generate clustering results, and calculate each cluster C k The embedding centroid of :
[0038] Select the sentence closest to the centroid from each cluster as the representative sentence of the cluster to form a streamlined data set;
[0039] S43, training model, select classification task as the fine-tuning task of BERT model, freeze some weights first, freeze the weights of BERT's Embedding layer and some Transformer layers in the initial stage, and only update the high-level weights and classification head to retain the general semantic characteristics of BERT;
[0040] S44. After the training is completed, the validation set is used to evaluate the results and obtain the accuracy of detecting fraudulent malicious apps.
[0041] The beneficial effects of the present invention are:
[0042] First, the present invention overcomes the limitation of directly using API names as features by refining API features. As a static label, the API name cannot fully reflect the contextual scenarios and calling patterns of its use. The fine-grained features proposed by the present invention, including the relationship between API parameters and APIs, the dependencies between APIs, calling frequency, calling sequence and combination patterns, and functional descriptions, greatly enrich the semantic expression of the API. This refined feature processing enables the model to capture complex behavior patterns more accurately and improves the ability to understand and analyze API usage.
[0043] Secondly, the present invention makes full use of the BERT model's sensitivity to language patterns by generating API feature sentences containing rich context. Compared with directly inputting refined feature characters, the generated feature sentences are more in line with natural language expression habits, which helps the BERT model to better play its advantage in understanding context. This processing method not only enhances the model's ability to capture fine-grained semantics, but also effectively eliminates the ambiguity that may be caused by API names and parameters, and improves the model's understanding and generalization capabilities.
[0044] Furthermore, the present invention proposes a method for generating proprietary semantic vectors in view of the limitations of fine-tuning the BERT model using API feature sentences directly. By constructing a corpus specific to the API field, training to obtain proprietary semantic word embeddings, and fusing them with BERT word embeddings and then fine-tuning them, the model can more deeply understand the uniqueness of the API proprietary domain language. This fused embedding space is closer to the actual distribution of the API feature corpus, which not only retains the contextual expression ability of the general language, but also enhances the proprietary semantic information of the API features, and improves the model's ability to accurately understand API behavior.
[0045] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0047] Figure 1 This is a schematic diagram of the overall process of the APP fraud detection method based on API semantic analysis of the present invention;
[0048] Figure 2 It is a detailed flow chart of the APP fraud detection method based on API semantic analysis of the present invention;
[0049] Figure 3 Schematic diagram of the fusion of the API-specific word embedding of the present invention and the BERT pre-trained word embedding. DETAILED DESCRIPTION
[0050] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0051] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0052] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0053] See also Figure 1 to Figure 3 , which is an APP fraud detection method based on API semantic analysis.
[0054] Example
[0055] This embodiment provides a detailed process of an APP fraud detection method based on API semantic analysis, and its flow chart is as follows: Figure 1 As shown, in general, it includes:
[0056] S1. Obtain APP samples, perform static analysis on the APP samples, decompile them to obtain the API code and structure of the APP samples, and obtain the parameter information in the API code; perform dynamic analysis on the APP, run the APP on the simulator, use debugging tools to monitor the APP behavior, and obtain the running log in the APP. The API code parameter information and API call information in the APP running log obtained above are the preliminary features of the APP, which will be further refined to obtain more sophisticated API features.
[0057] Build more sophisticated API features. By analyzing the API and its API parameters, we can obtain multi-dimensional features such as the association between APIs, API function descriptions, API call frequency, and API call order. Then, using these features and certain rules, we can generate API feature statements with natural language logic and detailed natural language descriptions of the refined features. These feature-refined API statements will serve as input for subsequent training of proprietary semantic word vectors and fine-tuning of the BERT model.
[0058] S2. Train proprietary semantic word vectors: Based on the API corpus and API feature sentences, use Word2Vec training to generate API proprietary semantic word vectors.
[0059] S3. Combine the context vector generated by BERT for weighted fusion to obtain the fused word vector, and embed the fused word vector into the word embedding layer of the BERT pre-trained model;
[0060] S4. Fine-tune the BERT model that incorporates proprietary semantic word vectors in the word embedding layer. Use API feature sentences to fine-tune the weights of the Embedding layer and Transformer layer of the BERT model. Use the clustering algorithm to extract representative sentences to optimize the model, input and detection efficiency. After training, input API feature sentences into the model to complete the prediction of fraud-related malicious apps.
[0061] like Figure 2 The detailed flow chart shown in FIG. 1 shows, step S1 specifically comprises the following steps:
[0062] S11. Obtain APP samples, obtain preliminary features through dynamic and static analysis, and perform detailed analysis on API features to finally obtain API feature statements;
[0063] S111. Perform static analysis on the APP sample, decompile it to obtain the API code and structure of the APP sample, and obtain the parameter information in the API code; and obtain the relationship between the API and other APIs by analyzing the API parameters to generate a feature statement of the relationship between the API and other APIs. First, analyze the input and output parameters of a certain API, extract the relationship between its parameters and other APIs, and then generate a feature statement according to the template-filled feature generation rules based on the relationship between the APIs. The rule for generating feature statements is to divide the feature statement into two parts: common features and differential features. The common feature is The parameter{parameter}of API{API1}contains the value obtained from API{API2}, and the differential feature is the filled part in the curly braces. The differential features are shown in Table 1:
[0064] Table 1
[0065]
[0066] The differences in the feature statements are shown in Table 1. For example, the analysis of API Calendar.CreateEvent() shows that the value of the parameter attendeeList comes directly from the data returned by API Contacts.Read(), and reading contact information has nothing to do with the function of creating calendar events. Then, a feature statement with a unified structure is generated based on the parameter information of the above APIs, and a preliminary structured statement is generated by filling in the differences. For example, the parameter attendeeList of API Calendar.CreateEvent() contains the value obtained from API Contacts.Read(), which is in English characters: The parameter attendeeList of API Calendar.CreateEvent contains the value obtained from APIContacts.Read.
[0067] S112. Obtain the API function description to generate a feature statement about the API function description. For the API function description, the acquisition method is to obtain the core function description of the API from the description field in the API document, and directly generate feature statements based on these function descriptions. For example, for Calendar.CreateEvent(), its function description is Calendar.CreateEvent() is used to create a new event entry in the user's calendar, supporting custom event times, locations and participant lists, and the English characters are Calendar.CreateEvent fith support for custom event times, locations and participant lists. For the API parameter description, the acquisition method is to parse the parameter description parameters part in the document to obtain the parameter description. For example, for the Calendar.CreateEvent(attendeeList) parameter, its description is The attendeeList parameter specifies the list of contacts that will participate in the event, supports multiple participants, and needs to be valid email addresses or user IDs, and the English characters are The attendeeList parameterspecifies the list of contacts that will participate in the event.Multipleparticipants are supported and need to be valid email addresses or user IDs.
[0068] S113. Analyze the calling frequency of the API to generate feature statements about the calling frequency of the API. By analyzing the application running log or API call stack, count the number of calls and calling frequency of each API in different scenarios, mark the APIs with abnormally high or low frequency calls, and generate feature statements according to certain rules based on the calling frequency. For example, the application running log shows that the API File.Read() is called 15,000 times / hour, while the calling frequency of the API in normal scenarios is usually no more than 300 times / hour. Then, based on the frequency information of the above APIs, generate feature statements of unified structure, generate preliminary structured statements by feature filling, and then perform semantic enhancement. Use the pre-trained language model to complete the generated statements semantically to ensure that the natural language expression is fluent and the logic is clear.
[0069] S114, analyzing the calling sequence of the API to generate feature statements about the API calls. Based on the calling relationship graph of the application, record the dependencies and calling paths of the API calls, form the calling sequence information reflecting the logical process, capture the logical features of the malicious behavior, and construct a malicious calling relationship graph, where each node represents an API call, each edge represents the calling sequence relationship, extract the API calling path and calculate the calling frequency, and for the calling path with high calling frequency, encode the calling path into a serialized text description, and generate feature statements according to certain rules based on the calling relationship. For example, in a malicious application, the calling sequence record shows: API A (network connection) → API B (obtaining geographic location information) → API C (uploading data), and then generate a feature statement of a unified structure based on the calling sequence information of the above APIs, generate a preliminary structured statement by filling in the feature, and then perform semantic enhancement, use the pre-trained language model, and complete the generated statement semantically to ensure that the natural language expression is fluent and the logic is clear.
[0070] S12. Preliminarily build an API corpus by collecting API call logs and API documents of applications. First, extract the key fields of API calls from the call logs, including API name, parameter information, return value, etc., parse the content of the API document, and extract information such as API function description, parameter description, and applicable scenarios; then clean the text information, remove special symbols and unify the capitalization; finally, segment the processed text to obtain ApiLex, a corpus containing rich API terms. The corpus is used in the segmentation process of generating proprietary semantic word vectors.
[0071] S2. Generate proprietary semantic word vectors. First, based on the word segmentation rules suitable for the API field, use the corpus to segment the refined API sentences, train the Word2Vec model, mine the word association relationship in the API call context in the corpus, and generate API proprietary semantic word vectors. Specifically, using the Skip-gram algorithm to train the Word2Vec model is an unsupervised training method. Before training, the setting parameters are adjusted to set the window size that determines the "window" range of the model when capturing the word context, set the word vector dimension to determine the dimension of the vector representation of each word, and set the minimum word frequency threshold. For words with a frequency lower than the threshold, they are ignored to reduce noise. During the training process, for a text sample, the model first initializes the vectors for the words in the text sample and assigns a randomly initialized vector to each word. Then, in each training iteration, the context sliding window slides on the text sample. During the sliding process, each word in the text sample will be used as the target word in the window. The model will generate the upper and lower words of the target word according to the window size. The model will calculate the loss function based on the gap between the generated upper and lower words and the actual upper and lower words. After multiple iterations, when the loss function converges, the training is completed and the API-specific semantic word vector is obtained.
[0072] S3, weighted fusion API proprietary word embedding and BERT pre-trained word embedding, such as Figure 3 For example, the API-specific word vectors generated by Word2Vec for each word are 768-dimensional, and the word vectors pre-trained by BERT are 728-dimensional. It is necessary to first use dimensionality reduction technology to reduce the higher-dimensional word vectors of the API-specific word vectors generated by Word2Vec and the word vectors pre-trained by BERT to align the word vector dimensions of the two models, and then use the weighted fusion strategy to assign weights to the two types of vectors. The weights are optimized according to specific tasks, and the vectors of the two models are added according to the weights. V f =αV w +βV b Finally, after obtaining the fused word vector, the word vector is embedded into the BERT pre-trained model.
[0073] S4. Fine-tune the BERT model.
[0074] S41. Improve the data set. The APP feature statements of the sample APP will be used as model input. In addition, the APP samples will be labeled according to whether the APP is involved in fraud to obtain a complete training data set.
[0075] S42, optimize the data set, cluster multiple feature sentences of each sample, obtain multiple representative feature sentences, and use the DBSCAN clustering method to cluster the API sentence embedding. First, obtain the API sentence of the existing weighted fusion model, E = {E 1 ,E2 ,…,E N}, select cosine similarity as the similarity measure in the clustering process, configure DBSCAN hyperparameters, embed the sentence E as the input of DBSCAN, perform density calculation and generate clustering results, and calculate each cluster C k The embedding centroid of : The sentence closest to the centroid is selected from each cluster as the representative sentence of the cluster to form a streamlined data set.
[0076] S43, training model, select classification task as the fine-tuning task of BERT model, freeze some weights first, freeze the weights of BERT's Embedding layer and some Transformer layers in the initial stage, and only update the high-level weights and classification head to retain the general semantic characteristics of BERT. After the training is completed, use the validation set to evaluate the results and obtain the accuracy of fraud-related malicious APP detection.
[0077] (1) Refining API features. Directly using API names as features has great flaws. The API name is only a static label and cannot reflect its actual context and calling mode. The same API may represent different functions in different scenarios, but the name itself cannot distinguish them. Its semantic expression is limited. The API name cannot provide semantic information such as calling logic and parameter configuration, and its semantic expression ability is weak. It is difficult to capture potential associations. The dependencies and combined calling modes between APIs cannot be reflected by the name, which makes it easy to miss the key logical chain in malicious behavior.
[0078] To address this problem, compared to directly using API names as features, using features that include the relationship between API parameters and APIs, the relationship between APIs, API call frequency, API call sequence and combination mode, and API and parameter function description as fine-grained features has many advantages. First, the relationship between APIs, the relationship between API parameters and APIs, and the API call sequence and combination mode have rich contextual information and more comprehensive semantic expressions, which can capture complex behavior patterns.
[0079] (2) Generate API feature sentences. Compared with directly inputting refined feature characters into the input model, generating API feature sentences containing rich context can better utilize BERT's sensitivity to language patterns and the BERT model's advantage in understanding context in natural language. At the same time, it has obvious advantages in capturing fine-grained semantics, eliminating ambiguity, enhancing the model's understanding ability, and improving the model's generalization ability.
[0080] Since the BERT model is very sensitive to contextual dependencies in natural language, it performs better when processing complete sentences. By generating sentences with rich contextual semantic structures, the model can better utilize BERT's language understanding capabilities to capture the functional details of the API. In contrast, it is difficult for the BERT model to fully utilize its advantages in understanding context by only inputting names and parameters.
[0081] Constructing API feature statements that include context can better capture fine-grained semantics. Although API names and parameters can express some semantics, they are usually relatively abstract and difficult to accurately reflect the contextual associations between APIs. For example, for APIs with similar names such as "getUserInfo" and "getAccountInfo", if the model only relies on name and parameter information, it may be difficult to distinguish their specific functional differences. Generating feature statements and describing their actual application scenarios (such as "obtaining user personal information" and "obtaining account information") helps the model understand the true meaning of the API more accurately.
[0082] Constructing API feature sentences containing context can better improve the generalization ability of the model. Generating sentences provides more natural language expressions and helps the model learn the usage of different APIs in various scenarios. For example, if the API is described as "querying the user's account balance" or "obtaining account information within a specified date" in English, the model can infer the functional meaning of other similar APIs, thereby enhancing the generalization ability.
[0083] Constructing API feature statements that include context can better eliminate ambiguity and enhance the model's understanding ability. API names often contain abstract expressions such as abbreviations and technical terms, which are prone to ambiguity. By describing the functions of the API in natural language through feature statements, more background information can be added to eliminate ambiguity. For example, the "sendData" API may be used in different scenarios (data transmission, message sending, etc.), and after generating the statement, it can be described as "sending data to the server" or "transmitting data in the network", so that the model has a more accurate understanding of the API function.
[0084] (3) Generating proprietary semantic vectors. Directly fine-tuning the BERT model using API feature sentences is simple, but has limitations and defects. First, features with API semantic uniqueness cannot be captured. In other words, directly fine-tuning the BERT model using API semantic sentences, the BERT model is a natural language model. The word embedding of BERT pre-training is mainly based on general natural language data, and is not optimized for the specific distribution of API feature sentence corpus. It cannot deeply understand the uniqueness of API-specific domain language in downstream tasks and cannot fit with API professional terminology. It may not be able to accurately capture the professional semantics of API features. In addition, API feature sentences contain technical terms such as call frequency, parameter-API relationship, and call sequence. This information may not fully reflect the difference in BERT's native embedding space, which leads to insufficient expression of specific semantic context when directly fine-tuning the BERT model using API feature sentences. Although directly fine-tuning BERT can adjust the weights, it cannot optimize the expression of API features from the bottom layer, which may lead to problems such as slow model training convergence and insufficient effect.
[0085] By comparison, generating API-specific semantic word embeddings and fusing them with BERT word embeddings before fine-tuning has many advantages. First, because permission information contains many unconventional English terms, constructing a corpus unique to the API field solves the problem that natural language models have difficulty understanding proprietary vocabulary. In addition, the fusion model focuses on the semantic expression of API features. The proprietary semantic word embeddings are trained based on the API feature sentence corpus, and can focus on key information in API features (such as API associations, call logic), providing a more refined feature representation. The fusion model also enriches BERT's word vector representation capabilities. Fusing proprietary semantic word embeddings with BERT native word embeddings can combine general semantic knowledge with API feature knowledge in specific fields, which not only retains the expressive power of the general language context, but also enhances the proprietary semantic information of API features. This fused embedding space is closer to the actual distribution of the API feature corpus, allowing the model to understand API behavior more accurately. In addition, the fusion model also enhances the feature generalization ability. The proprietary word embedding can capture the commonalities of general API behaviors through the training of API feature corpus, and the BERT pre-trained word embedding supplements the contextual semantics at the language level. This dual semantic expression enables the fused model to generalize new API features more strongly.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A method for detecting APP fraud based on API semantic analysis, characterized by: The method comprises the following steps: S1. Obtain APP samples, and perform static and dynamic analysis on the APP samples to obtain parameter information and operation logs of the APP samples, and perform refined feature analysis on the parameter information of the APP samples to construct API feature statements, and build an API corpus based on the APP sample operation logs; S2. Based on the API corpus and API feature sentences, use Word2Vec training to generate API-specific semantic word vectors; S3. Combine the context vector generated by BERT for weighted fusion to obtain the fused word vector, and embed the fused word vector into the word embedding layer of the BERT pre-trained model; S4. Fine-tune the BERT model in the word embedding layer. Use API feature statements to fine-tune the weights of the Embedding layer and Transformer layer of the BERT model. After the training is completed, predict the fraud-related malicious apps.
2. According to claim 1, a method for detecting APP fraud based on API semantic analysis is characterized by: In step S1, it includes: S11. Obtain APP samples, obtain preliminary features through dynamic and static analysis, and perform detailed analysis on API features to finally obtain API feature statements; S12. Preliminarily build an API corpus by collecting the application's API call logs and API documents.
3. According to claim 2, a method for detecting APP fraud based on API semantic analysis is characterized by: In step S11, by analyzing the API and its API parameters, and the APP behavior monitoring log, multi-dimensional features are obtained, including: the association relationship between APIs, API function description, API call frequency, and API call sequence; then the multi-dimensional features are used to generate rules for template-filled feature generation to generate API feature statements with natural language logic and natural language descriptions of refined features; refined feature generation includes: By analyzing the parameters of an API, the association relationship between the API and other APIs is obtained, and a feature statement of the association relationship between the API and other APIs is generated; wherein, the input and output parameters of a certain API are first analyzed, and the association relationship between its parameters and other APIs is extracted, and a feature statement is generated according to the rules of template filling feature generation through the association relationship between APIs; Obtaining API function descriptions and generating feature statements related to the API function descriptions; wherein, obtaining the core function descriptions of the API from the description field in the API document and directly generating feature statements based on the function descriptions; Analyze the calling frequency of APIs and generate feature statements about the calling frequency of APIs; by analyzing the application running log or API calling stack, count the number of calls and calling frequency of each API in different scenarios, mark APIs with abnormally high or low frequency calls, and generate feature statements according to the calling frequency according to the rules of template filling feature generation; Analyze the calling sequence of APIs and generate characteristic statements about API calls. Based on the calling relationship graph of the application, record the dependencies and calling paths of API calls to form calling sequence information that reflects the logical process, so as to capture the logical characteristics of malicious behavior and construct a malicious calling relationship graph, in which each node represents an API call and each edge represents the calling sequence relationship. Extract the API calling path and calculate the calling frequency. For the calling path with high calling frequency, encode the calling path into a serialized text description and generate characteristic statements according to the calling relationship.
4. According to claim 2, a method for detecting APP fraud based on API semantic analysis is characterized by: In step S12, the process of building the API corpus is: First, extract the key fields of the API call from the call log, including the API name, parameter information, and return value; Then, parse the API document content to extract the API function description, parameter description, and applicable scenario information; Then clean the text information, remove special symbols and unify the capitalization; Finally, the processed text is segmented to obtain the corpus ApiLex containing API terms.
5. According to claim 2, the method for detecting APP fraud based on API semantic analysis is characterized by: In step S2, the process of using Word2Vec training to generate API-specific semantic word vectors is as follows: Use the unsupervised training method Skip-gram algorithm to train the Word2Vec model; Adjust the settings before training. Set the window size that determines the range of the "window" when the model captures the context of the word, set the word vector dimension to determine the dimension of the vector representation of each word, and set the minimum word frequency threshold to ignore words with a frequency lower than the minimum word frequency threshold. During the training process, for a text sample, the model first initializes vectors for the words in the text sample and assigns a randomly initialized vector to each word. Then, in each training iteration, the context sliding window slides on the text sample. During the sliding process, each word in the text sample is used as the target word in the window, and the upper and lower words of the target word are generated according to the window size. The loss function is calculated based on the gap between the generated upper and lower words and the actual upper and lower words. After multiple iterations, when the loss function converges, the training is completed. After training is completed, API-specific semantic word vectors are obtained based on the API corpus and API feature sentences.
6. According to claim 5, a method for detecting APP fraud based on API semantic analysis is characterized in that: In step S3, the API-specific word vectors generated by Word2Vec and the word vectors pre-trained by BERT are first aligned, and the high-dimensional vectors are reduced in dimension using dimensionality reduction technology. Then, a weighted fusion strategy is used to assign weights to the two types of vectors. The weights are optimized according to specific tasks to obtain the fused word vectors: V f =αV w +βV b Where V f Represents the fused word vector, V w Represents the word vector generated by the Word2Vec model, V b Represents the word vector pre-trained by the BERT model.
7. According to claim 6, a method for detecting APP fraud based on API semantic analysis is characterized by: Embedding the fused word vector into the word embedding layer of the BERT pre-trained model means setting the fused word vector as the subword token of each input in the Embedding layer of the BERT pre-trained model.
8. According to claim 7, a method for detecting APP fraud based on API semantic analysis is characterized in that: In step S4, the fine-tuning training for the BERT model includes the following steps: S41. Improve the data set. The APP feature statements of the sample APPs will be used as model inputs. The APP samples will be labeled according to whether the APPs are involved in fraud, and a complete training data set will be obtained. S42. Optimize the data set, cluster multiple feature sentences of each sample, obtain multiple representative feature sentences, and use the DBSCAN clustering method to cluster the API sentence embeddings; First, obtain the API statement of the existing weighted fusion model, E = {E1, E2, ..., E N }, select cosine similarity as the similarity measure in the clustering process, configure DBSCAN hyperparameters, embed the sentence E as the input of DBSCAN, perform density calculation and generate clustering results, and calculate each cluster C k The embedding centroid of : Select the sentence closest to the centroid from each cluster as the representative sentence of the cluster to form a streamlined data set; S43, training model, select classification task as the fine-tuning task of BERT model, freeze some weights first, freeze the weights of BERT's Embedding layer and some Transformer layers in the initial stage, and only update the high-level weights and classification head to retain the general semantic characteristics of BERT; S44. After the training is completed, the validation set is used to evaluate the results and obtain the accuracy of detecting fraudulent malicious apps.