Method and System for Identifying Illegal Mobile Application Organizations Based on Multi-Level Feature Collaboration and Key Clue Tracing
Through multi-level feature collaboration and key clue traceability methods, combined with APK file analysis, dynamic sandbox operation and deep learning, the operational organization behind illegal mobile applications is identified, and the problem of low identification efficiency in the existing technology is solved, and efficient and accurate detection of illegal application is achieved.
Patent Information
- Application Number
- CN202510035037.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The existing illegal mobile application detection methods are difficult to effectively identify the operational organization behind illegal applications. There are problems such as low identification efficiency and insufficient correlation analysis, and it is impossible to fully cover the overall picture of their illegal behavior.
Through multi-level feature collaboration and key clue traceability methods, APK file analysis tool is used to extract key static information, combine dynamic sandbox operation to capture dynamic behavior and network traffic analysis, use deep learning models to generate high-dimensional features, and perform multi-level feature correlation analysis through graph neural networks to build the network topology of illegal organizations and dynamically update through reinforcement learning mechanisms.
It realizes rapid and efficient identification of illegal applications, reveals its operating model and risk characteristics, improves detection efficiency and accuracy, and adapts to changing illegal behavior patterns.
Smart Images

Figure CN119814461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a method and system for identifying illegal mobile application organizations based on multi-level feature collaboration and key clue tracing. Background Art
[0002] With the rapid development of mobile Internet, illegal applications (such as pornography, gambling, fraud, etc.) have emerged in an endless stream, posing a serious threat to user privacy, information security and social order. These illegal applications usually use complex technical means to evade the supervision of traditional detection methods by using a variety of development frameworks (such as Flutter, ReactNative, etc.), abusing permissions, forging signatures, and hiding network communications. These applications are not only concealed in their functional design, but also highly dynamic and decentralized during operation, making it difficult for a single detection dimension to fully cover the full picture of their illegal behavior. In addition, some illegal applications extend their functions by embedding third-party SDKs (such as advertising, payment, social, etc.), further increasing the difficulty of tracing and identification.
[0003] At present, the detection methods for illegal applications are mostly focused on single-dimensional analysis of permission analysis, signature feature comparison or network traffic. However, when faced with the organized operation of illegal applications, these methods have problems such as low recognition efficiency and insufficient correlation analysis, and cannot effectively explore the characteristics of the production and operation organizations behind illegal applications. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for identifying illegal mobile application organizations based on multi-level feature collaboration and key clue tracing, which can integrate the application's permission information, signature information, network traffic characteristics, dynamic behavior and log data, and through deep correlation analysis and clustering technology, quickly and efficiently identify the potential organizations behind illegal applications and reveal their operating models and risk characteristics.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] In the first aspect, a method for identifying illegal mobile application organizations based on multi-level feature collaboration and key clue tracing includes:
[0007] Step 1: Use APK file parsing tools to extract key static information from the code, including signatures, certificates, and permission configurations;
[0008] Step 2: Run the application through a dynamic sandbox to capture dynamic behaviors, including system calls and permission requests, and combine network traffic analysis to obtain the network communication characteristics of the application;
[0009] Step 3: Trace the key leads based on the key static information and the network communication characteristics of the application, and conduct traceability using the WHOIS query tool to obtain the traceability results, where the key leads include IP addresses and domain names;
[0010] Step 4: Use the deep learning model to perform feature learning based on the key static information, the network communication characteristics of the application, and the traceability results to generate high-dimensional features;
[0011] Step 5: Conduct multi-level feature correlation analysis using the graph neural network based on the high-dimensional features and the traceability results to obtain the hierarchical structure and behavior patterns within the organization;
[0012] Step 6: Based on the hierarchical structure and behavior patterns within the organization, construct the network topology of the illegal organization and dynamically update it through the reinforcement learning mechanism to adapt to the continuously changing illegal behavior patterns.
[0013] Furthermore, use the APK file parsing tool to extract the key static information from the code, including signatures, certificates, and permission configurations, including:
[0014] Extract code layer features using the APK file parsing tool;
[0015] Analyze the applications with reused signatures or forged signatures in combination with the historical database, mark the potential associated objects, and extract the distribution and dependency information of the third-party SDKs.
[0016] Furthermore, trace the key leads based on the key static information and the network communication characteristics of the application, and conduct traceability using the WHOIS query tool to obtain the traceability results, where the key leads include IP addresses and domain names, including:
[0017] Conduct traceability analysis on the key leads through the key static information and the network communication characteristics of the application;
[0018] Combine WHOIS queries, DNS resolution records, and the payment chain tracking system to extract the historical activity trajectories, server distributions, and fund flows of the domain name registrants;
[0019] Through the email addresses and contact information embedded in the resource files, conduct reverse lookups in combination with the public database to obtain the potential associations between illegal applications, where the potential associations between illegal applications are the traceability results.
[0020] Furthermore, use the deep learning model to perform feature learning based on the key static information, the network communication characteristics of the application, and the traceability results to generate high-dimensional features, including:
[0021] Use the hybrid deep learning model to perform feature learning and modeling based on the key static information, the network communication characteristics of the application, and the traceability results;
[0022] Pre-train the basic features through a Deep Belief Network (DBN) to extract the potential distribution relationship between the code layer and the traffic layer;
[0023] Use a Stacked Autoencoder (SAE) to fuse the features of the code layer and the traffic layer to generate high-dimensional features with rich semantics.
[0024] Furthermore, based on the high-dimensional features and the tracing results, use a graph neural network for multi-level feature correlation analysis to obtain the hierarchical structure and behavior patterns within the organization, including:
[0025] Conduct multi-level correlation analysis on the data of the code layer and the traffic layer according to the high-dimensional features and the tracing results;
[0026] Construct the causal relationship between features through causal inference technology, and combine with a graph neural network (GNN) to analyze the member roles and behavior patterns within the organization to obtain the results of feature correlation analysis.
[0027] Furthermore, the code layer features include signature information, certificate fields, permission configurations, component call relationships, control flow graphs, embedded URLs, and key data.
[0028] Furthermore, the key clues include IP addresses, domain names, payment accounts, email addresses, and contact information.
[0029] In a second aspect, an illegal mobile application organization recognition system based on multi-level feature collaboration and key clue tracing includes:
[0030] An extraction module for extracting key static information from the code using an APK file parsing tool, including signatures, certificates, and permission configurations;
[0031] An acquisition module for running the application through a dynamic sandbox to capture dynamic behaviors, where the dynamic behaviors include system calls and permission requests, and combining network traffic analysis to obtain the network communication characteristics of the application;
[0032] A tracing module for tracing the key clues based on the key static information and the network communication characteristics of the application, and using a WHOIS query tool for tracing to obtain the tracing results, where the key clues include IP addresses and domain names;
[0033] A learning module for performing feature learning using a deep learning model according to the key static information, the network communication characteristics of the application, and the tracing results to generate high-dimensional features;
[0034] An association analysis module for performing multi-level feature association analysis using a graph neural network according to the high-dimensional features and the tracing results to obtain the hierarchical structure and behavior patterns within the organization;
[0035] A building block for constructing the network topology of an illegal organization based on the hierarchical structure and behavior patterns within an organization, and dynamically updating it through a reinforcement learning mechanism to adapt to changing illegal behavior patterns.
[0036] In a third aspect, a computing device includes:
[0037] One or more processors;
[0038] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the described method.
[0039] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the described method.
[0040] The above solution of the present invention has at least the following beneficial effects:
[0041] Through the deep integration of code layer information mining and dynamic behavior capture technologies, combined with a hybrid deep learning model (stacked autoencoder and deep belief network), causal inference, and graph neural network analysis, accurate identification of the production and operation organizations of illegal applications (such as pornographic, gambling, and fraud APPs) is achieved, revealing their behavior patterns and organizational structures, and comprehensively improving the efficiency and accuracy of detecting illegal application organizations. Brief Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the overall functional structure of the implementation of the present invention. Detailed Embodiment
[0043] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Instead, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0044] As Figure 1 shown, an embodiment of the present invention proposes a method for identifying illegal mobile application organizations based on multi-level feature collaboration and key clue tracing, and the method includes the following steps:
[0045] Step 1: Use an APK file parsing tool to extract key static information from the code, including signatures, certificates, and permission configurations, specifically including:
[0046] Step 11: Use an APK file parsing tool (such as JADX) to decompile the target application and extract the static code structure, including information such as classes, methods, and fields. By parsing the Manifest file, extract the configuration information and permission declarations of components (such as Activity, Service, BroadcastReceiver), and analyze whether there are high-risk operations (such as dynamically loading code). Let the set of parsed static code be:
[0047]
[0048] Among them, represents the class and method information of the target application.
[0049] Preferably, a batch parsing tool can be combined to achieve the static code parsing of a large number of APK files and extract potential malicious behavior features.
[0050] Step 12: Extract the signature and digital certificate information from the APK file, including the issuer, subject, and validity period. Calculate the signature fingerprint through a hash value (such as SHA-256) and check whether it is consistent with the signatures of historical malicious applications. Extract payment information (such as Alipay, WeChat payment interfaces) and embedded payment accounts, and analyze the associated fund flows and security.
[0051] The extracted data set is:
[0052]
[0053] Among them, represents the signature hash, certificate content, and payment account information.
[0054] Preferably, a payment chain tracking system can be combined to locate the fund flows and the sources of funds for illegal organization operations.
[0055] Step 13: Extract information such as embedded email addresses, URLs, and domain names from the APK resource files, and combine domain name resolution records and the WHOIS database to analyze the domain name registrant information and historical activity trajectories. The set of extracted domain names and email information is:
[0056]
[0057] Among them, includes email addresses, domain names, and associated registrant information.
[0058] Preferably, the historical activity patterns of illegal domain names can be traced through a domain name attribution database (such as RDAP or WHOIS).
[0059] Step 14: Parse the list of permissions declared in the Manifest file, and focus on analyzing high-risk permissions (such as READ_SMS, ACCESS_FINE_LOCATION, etc.) and their distribution. Extract component configurations (such as the dynamic registration behavior of Broadcast Receiver), and mark unauthorized high-privilege calls.
[0060] The set of extracted permissions and component features is as follows:
[0061]
[0062] Among them, represents the permission types and call situations.
[0063] Preferably, the potential risks can be automatically marked in combination with the behavior rule library for high-risk permission calls.
[0064] Step 15: Extract sensitive information such as embedded URLs, encryption keys, and configuration files in the APK, and analyze whether these data point to known malicious domains or servers. Combine the embedded payment information to locate the illegal fund collection address. The set of extracted resource features is as follows:
[0065]
[0066] Among them, represents the hard-coded URL or key information.
[0067] Preferably, the legitimacy of sensitive information can be checked in combination with the domain name ownership database and the payment account verification tool.
[0068] Step 16: Detect the third-party SDKs used in the target application, including advertising SDKs, payment SDKs, social SDKs, etc., and extract their versions, call behaviors, and permission requirements. Analyze whether there are risks of privacy data collection or permission abuse in the SDK. The set of extracted SDK features is as follows:
[0069]
[0070] Among them, includes the SDK type and version information.
[0071] Preferably, the high-risk dependencies in illegal applications can be quickly identified in combination with the known high-risk SDK database.
[0072] Step 17: Use a static analysis tool to generate a call graph (CG) and a control flow graph (CFG) for the target application. Analyze the call paths of sensitive APIs (such as network requests and file operations), and mark potential malicious behavior paths. The extracted call relationships and control flow sets are as follows:
[0073]
[0074] Among them, represents the call path and control flow characteristics.
[0075] Preferably, specific sensitive call chains (such as dynamically loading DEX files) can be marked through a rule library.
[0076] Step 2: Run the application through a dynamic sandbox to capture dynamic behaviors. The dynamic behaviors include system calls and permission requests, and combined with network traffic analysis, obtain the network communication characteristics of the application, which may include:
[0077] Step 21: Drive the target application to run in a custom sandbox environment through an automated test framework (such as Appium or MonkeyRunner), and simulate regular user operation behaviors (such as login, browsing, clicking, etc.). The captured dynamic behaviors include: system call chains, dynamic permission requests, component loading processes (such as dynamically bound Services or dynamically registered BroadcastReceivers), and file operation behaviors (such as reading / writing files, dynamically loading resources, etc.). Let the set of captured dynamic behaviors be:
[0078]
[0079] Among them, represents the system calls, dynamic permissions, or file operation behaviors recorded during the running of the application.
[0080] Preferably, the sandbox environment can integrate an automated debugging tool (such as strace or ptrace) to accurately record the function stack of dynamic call behaviors.
[0081] Step 22: Deploy a traffic capture tool (such as Wireshark or Tcpdump) to collect the network communication traffic generated by the target application during operation in real time. The collected traffic includes DNS requests, HTTP / HTTPS communication packets, encrypted communication data, and traffic header information, etc. The set of collected raw traffic data is represented as:
[0082]
[0083] Among them, Represents information of a single data packet.
[0084] The traffic collection process can be modeled as:
[0085]
[0086] Where is the network communication behavior of the target application, is the captured time window.
[0087] Step 23: Parse and extract features from the collected traffic data, mainly including the following aspects:
[0088] (1) Communication frequency: Count the access frequency and time interval of each IP, and extract the short-term high-frequency access features:
[0089] ;
[0090] Where is the number of communications within a specific time period and represents the communication frequency feature.
[0091] (2) Data packet features: Analyze the data packet size, header information, and protocol distribution, and extract the application communication mode:
[0092] ;
[0093] Where represents the communication packet size distribution under a specific protocol.
[0094] (3) DNS request features: Extract the domain name, TTL value, and resolution result of DNS queries, and record the behavior of standby domain names and dynamic domain name switching:
[0095] ;
[0096] Where represents the domain name, TTL, and resolution result of a single DNS request.
[0097] Preferably, the short-term and long-term traffic behaviors can be analyzed by combining the time window division method:
[0098]
[0099] Step 24: Map the captured dynamic behavior and traffic features to a time series matrix:
[0100] ;
[0101] Among them, the rows of the matrix represent dynamic behaviors, traffic characteristics, and DNS requests, and the columns represent time steps.
[0102] Preferably, local feature extraction of the time series is performed through a sliding window mechanism, and the formula is:
[0103] ;
[0104] Step 25, the clustering result set is:
[0105] ;
[0106] Among them, represents a type of behavior pattern, such as batch requests or high-frequency communication. Preferably, combined with an anomaly detection algorithm (such as Isolation Forest) to mark high-frequency abnormal behaviors:
[0107]
[0108] Step 26, combine the dynamic behavior characteristics and traffic characteristics to generate a unified high-dimensional vector:
[0109] ;
[0110] Among them, is a fusion function, represents the generated high-dimensional feature vector. Preferably, cross-domain feature alignment can be achieved through a Generative Adversarial Network (GAN):
[0111] ;
[0112] Final feature vector is used for subsequent model training and deep feature learning.
[0113] Preferably, the whole process of dynamic behavior capture and traffic analysis can be realized through an automated script, which supports batch analysis and real-time feature extraction, thereby improving the efficiency and accuracy of illegal application analysis.
[0114] Step 3, based on the key static information and the network communication characteristics of the application, trace the key clues, and use the WHOIS query tool for source tracing to obtain the source tracing result. Among them, the key clues include IP addresses and domain names, specifically including:
[0115] Step 31, use the URL extracted in Step 1 and the DNS request data captured in Step 2 to construct a domain name set:
[0116] ;
[0117] Among them, is the target domain name. Use the WHOIS query tool to extract the registration information of each domain name, including fields such as the registrant's name, email, registration time, registrar, etc.:
[0118] ;
[0119] Let the set of registration information for all domain names be:
[0120] ;
[0121] Combined with email reverse lookup technology, analyze the historical activity tracks of the registrant and extract the set of associated domain names:
[0122] ;
[0123] Among them, is all domain names associated with a certain registered email.
[0124] Preferably, combined with the historical WHOIS database, analyze the domain name change and bulk registration behaviors of the registrant, and establish a domain name association matrix:
[0125] ;
[0126] Among them, indicates that domain names and belong to the same registrant. The tracing of domain name registrant information provides the preliminary relationship of nodes and edges for the basis association of illegal organizations.
[0127] Step 32, extract the set of payment accounts of the target application from static and dynamic analyses:
[0128] ;
[0129] Among them, includes account information, payment platform, and transaction serial number. Through the payment interface query tool, trace the fund flow of the payment account and construct a payment association matrix:
[0130] ;
[0131] Among them, indicates that there is a fund flow association between payment accounts and . Use association rule mining technology to analyze the fund transfer pattern and extract high-frequency transaction paths:
[0132] ;
[0133] Among them, is the threshold of the fund transfer amount.
[0134] Step 33, extract the DNS request set of the target application from the traffic data captured by dynamic analysis:
[0135] ;
[0136] For each domain name 's resolution record, extract the set of IP addresses it resolves to:
[0137] ;
[0138] Combine historical resolution records and geolocation tools to analyze the geographical distribution and affiliated operators of IP addresses:
[0139] ;
[0140] Among them, represents the affiliated information of the IP address.
[0141] Utilize the association relationship between domain names and IP addresses to construct a DNS association matrix:
[0142] ;
[0143] Among them, indicates that the domain name and resolve to the same IP address.
[0144] Step 34, combine domain name registration information, payment accounts, and DNS resolution records to construct a preliminary association network:
[0145] ;
[0146] Among them, the node set IP · the edge set has an association
[0147] Calculate the importance for each node in the network, estimated by the PageRank algorithm:
[0148] ;
[0149] Among them, is the importance of node ; is the neighbor set of node ; is the degree of node ; is the damping factor (take 0.85).
[0150] Step 4: Based on the key static information, the applied network communication characteristics, and the traceability results, use a deep learning model for feature learning to generate high-dimensional features, specifically including:
[0151] Step 41: Design a hybrid deep learning framework composed of a deep belief network (DBN) and a stacked autoencoder (SAE). First, layer by layer, the DBN learns the basic distribution relationship of the code layer and traffic layer features, and then the SAE is used to mine the non-linear structure and optimize the feature fusion representation.
[0152] (1) Pre-training of the deep belief network (DBN):
[0153] Let the code layer feature set be , and the clue traceability feature be . The input feature is defined as:
[0154]
[0155] Through the restricted Boltzmann machine (RBM), the input feature is pre-trained layer by layer to learn the basic distribution and initialize the model parameters. The energy function of the RBM is defined as:
[0156] ;
[0157] where and are the units of the visible layer and the hidden layer respectively; is the connection weight; and are the bias terms.
[0158] By minimizing the energy function, the initial representation of the feature H DBN is learned.
[0159] (2) Non-linear modeling of the stacked autoencoder (SAE):
[0160] Take the output of the DBN as the input of the SAE, and perform non-linear compression representation of the features through layer-by-layer encoding and decoding to capture the high-order feature semantics. The encoding and decoding processes of the SAE are defined as:
[0161] ;
[0162] ;
[0163] where and represent the weights and biases of the th layer, and are the activation function and the decoding function respectively, is the hidden representation of the
[0164] Step 42: Perform deep modeling on the code layer features , traffic layer features and clue features .
[0165] (1) Code layer feature modeling:
[0166] Perform layer-by-layer training on the code layer features (such as signature information, permission configuration, SDK calls, etc.), extract the potential associations between features, and generate implicit representations:
[0167] ;
[0168] (2) Traffic layer feature modeling:
[0169] Perform non-linear compression modeling on the traffic layer behavior features (such as system calls, traffic patterns, dynamic loading paths, etc.) to generate high-order feature representations:
[0170] ;
[0171] (3) Clue feature modeling:
[0172] Perform modeling on the traceability features (such as domain names, payment accounts, DNS resolution records, etc.) to generate associated representations:
[0173] ;
[0174] (4) Fusion process:
[0175] Use the weights of the middle layer of the DBN to initialize the parameters of the SAE, further optimize the global representation after fusion, and generate high-dimensional semantic vectors:
[0176] ;
[0177] Among them, represents the feature fusion function.
[0178] Step 43, to adapt to the diverse behavioral patterns of illegal mobile applications, the present invention optimizes the hybrid model in combination with the transfer learning mechanism, enabling it to maintain high - efficiency analysis capabilities in different illegal behavior scenarios. The core of transfer learning lies in leveraging the knowledge of the source domain (such as existing illegal application feature data) to guide the generalization ability of the model in the target domain (such as new illegal behavior feature data). The loss function of transfer learning is defined as:
[0179] ;
[0180] where, represents the loss of the source domain (known illegal application dataset), which is used to maintain the stability of the model on existing data, represents the loss of the target domain (new illegal behavior dataset), which is used to guide the model to adapt to new features, is the weight coefficient, which is used to balance the learning of the two domains.
[0181] During the transfer process, the model first extracts the basic feature representation of the source domain through a deep belief network (DBN), and then uses a stacked auto - encoder (SAE) to perform a non - linear mapping on the features of the target domain. By minimizing the transfer loss function, the model can efficiently learn new illegal behavior features while retaining the expressive ability of existing features.
[0182] Step 44, after completing deep learning and transfer optimization, the hybrid model of the present invention outputs a high - dimensional semantic vector that fuses the feature of the code layer (code structure, signature information, permission configuration) and the feature of the traffic layer (system call chain, network traffic pattern) for subsequent causal inference and organizational association analysis. The output high - dimensional feature vector is represented as:
[0183] ;
[0184] where, represents the feature of a certain dimension, which reflects the complex association between the code layer and the dynamic behavior, The generation process of combines cross - domain feature alignment and deep compression representation of multi - level data, ensuring the semantic consistency and discrimination ability of the vector.
[0185] This high - dimensional semantic vector can not only reveal the behavioral patterns of illegal applications but also effectively associate the potential cooperation relationships among gang members. This high - dimensional semantic vector can reveal the behavioral patterns of illegal applications and their potential relationships with organizational members, providing key input data for subsequent causal inference and organizational analysis. The output feature vector will be directly used in Step 5 (multi - level feature association and organizational analysis) and Step 6 (construction of illegal organization network).
[0186] Step 5. Based on the high-dimensional features and the traceability results, use a graph neural network to perform multi-level feature correlation analysis to obtain the hierarchical structure and behavior patterns within the organization, specifically including:
[0187] Step 51. Analyze the features at the code layer , the features at the traffic layer and the traceability data to construct a causal relationship network of the illegal software organization . The node set includes entities such as applications, domain names, payment accounts, and server IPs. The set represents the causal associations between entities (e.g., a payment account supports the registration of a domain name ). Use the conditional independence test method (such as the PC algorithm) to infer causal relationships: ;
[0188] The causal strength is described by a structural equation model (SEM):
[0189] ;
[0190] where represents the causal coefficient, is the noise term.
[0191] Step 52. Combine the relationships in the causal network to construct a feature matrix containing the code layer, the traffic layer, and the traceability data:
[0192] ;
[0193] where represents the th feature of the th entity (such as the number of payment times, the number of DNS resolutions). Perform dimensionality reduction through principal component analysis (PCA):
[0194] ;
[0195] where is the dimensionality reduction weight matrix, is the feature representation after dimensionality reduction. The features after dimensionality reduction are used to identify the characteristics of different members in the illegal organization.
[0196] Step 53. Use the high-dimensional semantic vectors generated in Step 54 as node features, and combine them with the causal relationship network to perform organizational analysis and identification through a graph neural network (GNN).
[0197] (1) Node embedding update rule:
[0198] The node embedding update of the GNN is defined as:
[0199] ;
[0200] where, represents the feature embedding of node at the -th layer; is the neighbor set of node ; and are the weight and bias at the -th layer; is the non-linear activation function.
[0201] (2) Community division and clustering analysis:
[0202] The embedded nodes are divided into communities by the Louvain algorithm, with the goal of maximizing modularity
[0203] ;
[0204] where, is an element of the adjacency matrix, indicating whether there is an edge between nodes and ; and are the degrees of the nodes; is the Kronecker function, indicating whether nodes and belong to the same community.
[0205] (3) Organizational relationship and hierarchical structure analysis:
[0206] The set of community division results is:
[0207] ;
[0208] where, represents a subgroup of the illegal organization. Based on the division results, the core members of the organization and their collaboration relationships are analyzed to reveal the hierarchical structure and behavior patterns of the illegal software organization.
[0209] Step 6, based on the hierarchical structure and behavior patterns within the organization, construct the network topology of the illegal organization and dynamically update it through a reinforcement learning mechanism to adapt to the changing illegal behavior patterns, specifically including:
[0210] The construction and dynamic update of the organizational network topology are as follows:
[0211] Step 61, utilize the causal relationship network and the clustering results of the graph neural network (GNN) , construct a dynamic topology graph of illegal software organizations .
[0212] Node set includes members of illegal organizations (such as applications , domain names , payment accounts and server IPs ); The edge set represents the association relationship between members (such as capital flow, domain name reuse, code collaboration).
[0213] The edge weights are calculated by combining and weighting different relationships between members:
[0214] ;
[0215] Among them, is the capital flow intensity, measuring the frequency and amount of capital flow between the payment account and ; is the degree of domain name or IP address reuse, measuring whether the nodes and resolve to the same IP address; is the similarity of features in the code layer and traffic layer, calculated through the high-dimensional feature vector ; is a tuning parameter used to balance the weights of different relationships.
[0216] In the edge weight formula, is defined as:
[0217] ;
[0218] Among them, and are the high-dimensional semantic vectors of the nodes and .
[0219] is used to measure the hierarchy of the organization by calculating the modularity index of the topology graph
[0220] ;
[0221] Among them, represents the adjacency matrix element, indicating whether there is an edge between the nodes and ; is the degree of the node; total number of edges; Is the Kronecker function, representing nodes and whether they belong to the same community.
[0222] Step 62, dynamic update and adaptive optimization:
[0223] Based on Reinforcement Learning (RL) technology, dynamically update the topology graph to make it adapt to new types of illegal behaviors.
[0224] (1) State and action definitions:
[0225] The state represents the structure of the current topology graph, expressed as a set of nodes and edges . The action is an operation on the graph, such as adding a node , adding an edge , or adjusting the edge weight .
[0226] (2) Reward function definition:
[0227] The reward function measures the contribution of the added nodes and edges:
[0228] ;
[0229] where is the behavior coverage rate, representing the proportion of illegal nodes successfully detected by the model:
[0230] ;
[0231] The modularity exponent, which measures the tightness of the organizational hierarchy.
[0232] (3) Dynamic adjustment and optimization:
[0233] The optimization goal of reinforcement learning is to maximize the expected return of the policy
[0234] ;
[0235] where is the value function of the next state, is the discount factor, used to balance short-term and long-term rewards.
[0236] Optimize the policy through the policy gradient method, so that the topology graph can adapt to new illegal behavior patterns
[0237] (4)Node importance update:
[0238] Based on the update results of reinforcement learning, recalculate the node importance (such as PageRank);
[0239] ;
[0240] where, is the importance of node ; is the neighbor set of node ; is the degree of node ;
[0241] Step 63, finally output the dynamically updated illegal organization network topology , including the following contents: (1) Organizational hierarchy: showing the core members of the illegal software organization and their division of labor; (2) Real-time update structure: combining the adjustment of new nodes and edges to maintain the real-time and robustness of the model; (3) Behavior pattern coverage: through the dynamic update mechanism, efficiently capture new illegal behaviors. The above process is the illegal mobile application organization recognition technology based on multi-level feature collaboration and key clue tracing.
[0242] Figure 1 Among them, the illegal software database is used to store the APK files of illegal software and their associated data, including static code layer features, dynamic traffic layer features, domain name tracing information, payment chain data, and historical analysis results, providing comprehensive data support and model verification basis for the entire analysis process.
[0243] Static information mining and feature extraction are used to parse the APK file of the target application to extract signature information, certificate fields, permission configurations, component call relationships (CG), control flow graphs (CFG), and embedded URLs and keys, etc. code layer features, which are used to mark potential malicious behaviors and provide references for dynamic analysis.
[0244] Dynamic behavior capture and network traffic analysis are used to run the target application in a sandbox environment to capture the runtime system call chain, dynamic permission requests, component loading, and file operation behaviors, and combine network traffic capture tools to extract DNS requests, communication patterns, alternative domain names, and packet behaviors, etc. dynamic traffic layer features.
[0245] Code layer features are used to extract static code data such as signatures, certificates, permission configurations, and embedded URLs extracted from static information mining, as preliminary marked features of malicious behaviors, which are used for subsequent correlation analysis and modeling.
[0246] Flow layer features are used for runtime data obtained through dynamic behavior capture, including system call chains, permission usage patterns, network communication behaviors, and dynamic loading paths, etc., to reveal malicious behavior patterns and potential collaboration relationships.
[0247] Key clue tracing and source tracing are used to utilize clues such as domain names, payment accounts, email addresses, and IP addresses in code layer features and flow layer features, and through WHOIS queries, payment chain tracing, and DNS resolution history, locate the associated nodes of illegal organizations and construct a preliminary multi-dimensional association network.
[0248] Deep Belief Network (DBN) pre-training is used to pre-train multi-level features using a deep belief network, learn the basic distribution of static and dynamic data, and provide optimized weight initialization for subsequent deep non-linear modeling.
[0249] Stacked Autoencoder (SAE) multi-layer network mining is used to further fuse the non-linear relationship between the code layer and dynamic features through a stacked autoencoder, generate high-dimensional semantic feature representations, capture complex behavior patterns, and support organizational analysis and network topology modeling.
[0250] Multi-level feature association and organizational analysis are used to combine the results of deep feature learning and key clue tracing information, use causal inference techniques to construct a causal relationship network among organizational members, and analyze the internal structure of the organization through a Graph Neural Network (GNN) to reveal core members and collaborative behaviors.
[0251] Illegal organization network construction is used to construct a dynamic network topology map of illegal organizations based on the results of multi-level feature association and organizational analysis, display hierarchical relationships and behavior division of labor, and dynamically update the network model through reinforcement learning to adapt to new types of illegal behavior patterns, providing accurate decision-making support for supervision and governance.
[0252] The present invention innovatively combines code layer and flow layer feature analysis, and through multi-level data fusion and deep learning techniques, realizes the precise identification of the behaviors of illegal software organizations and clue tracing. Compared with traditional methods, the present invention shows higher robustness and adaptability when dealing with new types of illegal behaviors and complex association patterns.
[0253] By designing a hybrid deep learning framework, the present invention can mine hidden relationships of illegal software organizations from multi-level features such as signature information, permission configuration, payment data, network traffic, and dynamic behaviors, and reveal their collaboration patterns and resource sharing mechanisms. Causal inference and graph neural networks further enhance the ability to analyze the relationships among organizational members, providing strong support for association modeling.
[0254] The dynamic update and reinforcement learning mechanism enables the system to optimize the illegal organization network topology in real time, adapt to the evolving illegal behavior patterns, and improve the timeliness and accuracy of analysis. The present invention also provides the ability to quickly respond to new threats through high-dimensional semantic feature generation and dynamic feedback.
[0255] The present invention demonstrates wide applicability in the detection of illegal software organizations, and can also be extended to applications such as malware monitoring, botnet analysis, and the prevention and control of other complex network security threats, providing accurate and efficient technical support for regulatory and law enforcement agencies. It is an intelligent solution for complex network environments.
[0256] An illegal mobile application organization recognition system based on multi-level feature collaboration and key clue tracing includes:
[0257] An extraction module for extracting key static information from the code using an APK file parsing tool, including signatures, certificates, and permission configurations;
[0258] An acquisition module for running the application through a dynamic sandbox to capture dynamic behaviors, where the dynamic behaviors include system calls and permission requests, and combining network traffic analysis to obtain the network communication characteristics of the application;
[0259] A tracing module for tracing key clues based on the key static information and the network communication characteristics of the application, and using the WHOIS query tool for tracing to obtain the tracing results, where the key clues include IP addresses and domain names;
[0260] A learning module for performing feature learning using a deep learning model according to the key static information, the network communication characteristics of the application, and the tracing results to generate high-dimensional features;
[0261] An association analysis module for performing multi-level feature association analysis using a graph neural network according to the high-dimensional features and the tracing results to obtain the hierarchical structure and behavior patterns within the organization;
[0262] A construction module for constructing the network topology of the illegal organization based on the hierarchical structure and behavior patterns within the organization, and dynamically updating it through a reinforcement learning mechanism to adapt to the changing illegal behavior patterns.
[0263] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An illegal mobile application organization recognition method based on multi-level feature collaboration and key clue tracing, characterized in that The method includes: Step 1, using an APK file parsing tool to extract key static information from the code, including signatures, certificates, and permission configurations; Step 2, running the application through a dynamic sandbox to capture dynamic behaviors, where the dynamic behaviors include system calls and permission requests, and combining network traffic analysis to obtain the network communication characteristics of the application; Step 3, based on the key static information and the network communication characteristics of the application, tracking key clues and using the WHOIS query tool for traceability to obtain traceability results, where the key clues include IP addresses and domain names; Step 4, according to the key static information, the network communication characteristics of the application, and the traceability results, using a deep learning model for feature learning to generate high-dimensional features; Step 5, according to the high-dimensional features and the traceability results, using a graph neural network for multi-level feature correlation analysis to obtain the hierarchical structure and behavior patterns within the organization; Step 6, based on the hierarchical structure and behavior patterns within the organization, constructing the network topology of the illegal organization and dynamically updating it through a reinforcement learning mechanism to adapt to the constantly changing illegal behavior patterns.
2. The illegal mobile application organization recognition method based on multi-level feature collaboration and key clue tracing according to claim 1, wherein, Using an APK file parsing tool to extract key static information from the code, including signatures, certificates, and permission configurations, including: Using an APK file parsing tool to extract code layer features; Combining historical database analysis to analyze applications with reused signatures or forged signatures, marking potential associated objects, and extracting third-party SDK distribution and dependency information.
3. The method for identifying an illegal mobile application organization based on multi-level feature collaboration and key clue tracing according to claim 2, wherein Based on the key static information and the network communication characteristics of the application, tracking key clues and using the WHOIS query tool for traceability to obtain traceability results, where the key clues include IP addresses and domain names, including: Tracking and tracing key clues through the key static information and the network communication characteristics of the application; Combining WHOIS queries, DNS resolution records, and payment chain tracking systems to extract the historical activity trajectories, server distributions, and fund flows of domain name registrants; Through the emails and contact information embedded in the resource files, combined with public databases for reverse lookup to obtain potential associations between illegal applications, where the potential associations between illegal applications are traceability results.
4. The method for identifying an illegal mobile application organization based on multi-level feature collaboration and key clue tracing according to claim 3, wherein According to the key static information, the network communication characteristics of the application, and the traceability results, using a deep learning model for feature learning to generate high-dimensional features, including: According to the key static information, the network communication characteristics of the application, and the traceability results, using a hybrid deep learning model for feature learning and modeling; Pre-training the basic features through a deep belief network to extract the potential distribution relationships between the code layer and the traffic layer; Using a stacked autoencoder to fuse the code layer and traffic layer features to generate semantically rich high-dimensional features.
5. The method for identifying an illegal mobile application organization based on multi-level feature collaboration and key clue tracing according to claim 4, wherein According to the high-dimensional features and the traceability results, using a graph neural network for multi-level feature correlation analysis to obtain the hierarchical structure and behavior patterns within the organization, including: Performing multi-level correlation analysis on the data of the code layer and the traffic layer according to the high-dimensional features and the traceability results; Constructing the causal relationships between features through causal inference techniques, and combining with a graph neural network to analyze the member roles and behavior patterns within the organization to obtain the results of feature correlation analysis.
6. The method for identifying an illegal mobile application organization based on multi-level feature collaboration and key clue tracing according to claim 5, characterized in that The code layer features include signature information, certificate fields, permission configurations, component call relationships, control flow graphs, embedded URLs, and key data.
7. The illegal mobile application organization recognition method based on multi-level feature collaboration and key clue tracing according to claim 6, characterized in that, The key clues include IP addresses, domain names, payment accounts, email addresses, and contact information.
8. An illegal mobile application organization recognition system based on multi-level feature collaboration and key clue tracing, characterized in that, The system is used to execute the method described in any one of claims 1 to 7, and includes: An extraction module, configured to use an APK file parsing tool to extract key static information from the code, including signatures, certificates, and permission configurations; An acquisition module, configured to run the application through a dynamic sandbox to capture dynamic behaviors, where the dynamic behaviors include system calls and permission requests, and combine network traffic analysis to obtain the network communication characteristics of the application; A tracing module, configured to trace the key clues based on the key static information and the network communication characteristics of the application, and use the WHOIS query tool for tracing to obtain a tracing result, where the key clues include IP addresses and domain names; A learning module, configured to perform feature learning using a deep learning model according to the key static information, the network communication characteristics of the application, and the tracing result to generate high-dimensional features; An association analysis module, configured to perform multi-level feature association analysis using a graph neural network according to the high-dimensional features and the tracing result to obtain the hierarchical structure and behavior patterns within the organization; A construction module, configured to construct the network topology of the illegal organization based on the hierarchical structure and behavior patterns within the organization, and dynamically update it through a reinforcement learning mechanism to adapt to the continuously changing illegal behavior patterns.
9. A computing device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Deep tracing method for malicious sample
CN109361643A
Threat tracing method and related equipment
CN112131571A