Privacy risk measurement method for Android mobile application program

Through multi-dimensional evaluation and self-learning weight allocation mechanism, combined with static analysis and dynamic analysis, the inaccuracy problem of Android application privacy risk assessment in the existing technology is solved, and a comprehensive and accurate assessment of Android application privacy risks is achieved, which improves data security.

CN120296778APending Publication Date: 2025-07-11XIANGTAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510330238.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When evaluating the privacy risks of Android applications, the existing technology has problems such as single evaluation indicators, unreasonable selection of measurement indicators, and limitations of evaluation algorithms, resulting in inaccurate and incomplete evaluation results, and the inability to effectively identify potential privacy risks.

Method used

Multi-dimensional evaluation indicators are adopted, combined with static analysis and dynamic analysis, and through evaluations of permission requests, data sources, data flow directions, API communication security, etc., a risk matrix is generated and K-means clustering analysis is carried out, and a self-learning weight allocation mechanism is combined to achieve privacy risk assessment of Android applications.

Benefits of technology

It realizes a more comprehensive and accurate assessment of the privacy risks of Android applications, helping developers identify and prevent potential privacy leaks and improve data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296778A_ABST
    Figure CN120296778A_ABST
Patent Text Reader

Abstract

The invention relates to a privacy risk measurement method for an Android mobile application program, and aims to comprehensively evaluate privacy risks in the application program by combining static analysis and dynamic analysis. According to the method, a multi-dimensional risk assessment model is introduced, and the method comprises the steps of excessive permission request detection, data source verification, data flow terminal analysis, HTTP / HTTPS channel privacy leakage detection and secure transmission verification, and API communication security analysis. Through a self-learning weight distribution mechanism, the weight of risk factors can be adjusted, a weight coefficient for a function scene is generated, and data classification and risk assessment are performed in combination with a Mahalanobis distance and a K-means clustering algorithm. According to the method, a developer can be helped to identify the risk of privacy disclosure more accurately, the data security is improved, and potential privacy threats are prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical fields of privacy protection and data security for privacy protection of Android application programs, and particularly relates to a framework for a privacy risk sensitivity measurement method for privacy data in Android mobile application programs. Background Art

[0002] With the rapid development of the mobile Internet, smart phones have become an indispensable part of people's daily lives. As a mobile operating system with a high global market share, Android applications (hereinafter referred to as "Apps") collect a large amount of personal information of users while providing diverse functions. However, more and more Android application programs have risks of excessive permission requests, unauthorized data sharing, and sensitive information leakage. These privacy issues not only threaten the data security of users, but also may trigger malicious attacks and privacy violations.

[0003] At present, some privacy risk assessment methods rely on the permission analysis of applications and conduct privacy risk assessment on applications based on the sensitivity of permissions. There are also some methods that rely on the actual behavior of applications to evaluate privacy risks. With the popularity of mobile applications, the protection of users' privacy data has become particularly important. However, there are some problems in the measurement of privacy data sensitivity in the existing technologies, which pose potential risks and threats to users' privacy. Specifically, the main limitations of the existing methods include the following three aspects: 1) Singularity of evaluation indicators: Many existing methods only consider indicators in a single dimension, such as only considering the management of permission requests or the encryption of data transmission. This method cannot comprehensively cover various sources of privacy risks in mobile applications, and this evaluation method cannot comprehensively cover various sources of privacy risks in mobile applications, resulting in certain limitations in the evaluation results and making it difficult to accurately identify potential privacy risks, leading to poor privacy protection effects; 2) Irrationality of selected measurement indicators: Although some existing methods use multiple indicators, these indicators often cannot fully reflect the actual privacy risks and have little impact on privacy risks, and cannot accurately evaluate the actual privacy risks. For example, some methods only evaluate the transmission path of data while ignoring key factors such as the data source, the credibility of the terminal, and the security of third-party libraries; 3) Limitations of evaluation algorithms: In some privacy risk assessment methods, classification algorithms or rules are often used for risk assessment. These algorithms are difficult to handle complex data relationships and lack the mutual influence between multiple risk factors, resulting in limited accuracy of the evaluation results. Many existing methods use simple classification algorithms to evaluate privacy risks, and these algorithms often cannot handle complex data relationships, resulting in inaccurate privacy risk assessment results. For example, some methods classify mobile applications only based on simple rules or thresholds while ignoring the mutual influence between different indicators.

[0004] In summary, to meet the multi-dimensional requirements of privacy risk assessment and leveraging the advantages of advanced algorithms, the assessment scope is extended to multiple key dimensions such as permission requests, data flow, and API communication security. When traditional single assessment methods are difficult to accurately reflect the privacy risks of applications, more accurate assessment decisions can be made by selecting indicators that have a significant impact on privacy risks. In addition, for the risk characteristics of mobile applications, through complex data processing techniques such as calculating the Mahalanobis distance between risk characteristics and K-means based clustering analysis, the limitations in privacy risk assessment can be reduced, and the accuracy and dimensional integrity of privacy risk analysis can be optimized, thereby improving the overall efficiency of the assessment method. To solve the above problems, the following three aspects need to be addressed: 1) Multi-dimensional evaluation indicators: Multi-dimensional evaluation indicators should be adopted to comprehensively cover various sources of privacy risks in mobile applications. For example, not only the management of permissions should be considered, but also the sources of data, data flow, the behavior of third-party libraries, etc. should be evaluated; 2) Find relatively reasonable measurement indicators: Measurement indicators that have a greater impact on privacy risks should be selected to ensure that these indicators can accurately reflect the actual privacy risks. For example, evaluate whether the terminals of data flow are trustworthy and whether there is privacy leakage during data transmission; 3) Find relatively reasonable measurement methods: Reasonable classification methods should be adopted to evaluate privacy risks and be able to handle complex data relationships. For example, advanced algorithms such as fuzzy clustering analysis can be used to comprehensively evaluate the mutual influence between multiple indicators and accurately identify the privacy risks of mobile applications. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the existing measurement methods and propose a privacy risk measurement method framework for Android mobile applications. This method combines static analysis and dynamic analysis and can accurately assess the privacy risks of Android applications. Based on the static analysis framework for mobile applications, it can accurately measure the privacy risks of data in mobile applications to measure the privacy risks of mobile applications. This framework not only considers traditional risk factors such as over-authorization of application data permission requests, data source security, and data terminal security flow, but also covers network risk factors such as the security of data during transmission, thereby qualitatively evaluating the application program. In addition, this method also integrates a self-learning weight allocation mechanism that can automatically allocate the weights of 5 risk factors according to actual application behaviors and quantitatively evaluate the application program. This method can evaluate the privacy risks of Android applications with finer granularity and higher accuracy.

[0006] Step 1: Factor extraction:

[0007] The privacy risk measurement method for Android applications provided by the present invention selects 5 key indicators, which have been experimentally verified as essential components for comprehensively evaluating the privacy risks of mobile applications.

[0008] 1) Detection of excessive permission requests: Permission management is the core of the Android security architecture. Excessive permission requests may lead to unnecessary privacy risks. Therefore, evaluating whether an application requests permissions that exceed its functional requirements is an important criterion for measuring the privacy risk of the application. By using a static analysis tool to extract the permission declarations of the mobile application and analyzing whether the extracted permissions match the scenario requirements declared by the application, it is determined whether the permission application exceeds the minimum requirements in this scenario to evaluate potential privacy risks.

[0009] 2) Verification of data sources: Trustworthy data sources are an important foundation for ensuring the security of application programs. Unreliable data sources may become channels for malicious attacks and data contamination. Therefore, the analysis of data sources is crucial for ensuring the overall security of application programs. By using a static analysis tool to trace data source nodes, such as API calls or external inputs, the trustworthiness and security of the data sources are confirmed to evaluate whether the data sources may pose a privacy leakage risk.

[0010] 3) Analysis of data flow to the terminal: The final storage or processing location of data directly affects data security. If data is transmitted to an insecure or unverified terminal, even if the data source and transmission process are secure, it may still lead to privacy leakage. By using a static analysis tool to trace the data flow to the terminal and evaluating the security of the terminal, it is ensured that the data will not be transmitted to high-risk or untrusted terminals.

[0011] 4) Detection of privacy leakage and verification of secure transmission in HTTP / HTTPS channels: Encrypted data streams can hide the transmission of sensitive information, making it difficult for traditional analysis tools to detect. This increases privacy risks because some applications may use these encrypted channels to leak user data. Some mobile applications also use custom encryption and non-standard protocols for data transmission, which are usually not monitored by conventional security tools, increasing the difficulty of detection and defense. By using a static analysis tool, these potential privacy leakage paths can be identified to help developers better understand and prevent privacy risks.

[0012] 5) Security analysis of API communication: The security of API communication is crucial for ensuring the confidentiality and integrity of data during transmission. By evaluating the security of API communication, potential security vulnerabilities can be identified, and corresponding measures can be taken to prevent data from being intercepted or tampered with by attackers during transmission.

[0013] Step 2: Generate a risk matrix:

[0014] 1) Data preprocessing: Process the collected relevant information, and convert the information extracted from each privacy risk factor into a one-dimensional vector form for further analysis;

[0015] 2) Generate a risk feature matrix: Construct a privacy risk feature matrix with a specification of 5*4, which is used to represent the data information of each risk factor;

[0016] 3) Matrix dimensionality reduction: Perform dimensionality reduction on the generated risk feature matrix to obtain its representation form in a two-dimensional space, facilitating subsequent clustering analysis.

[0017] Step Three: K-means clustering analysis

[0018] 1) Select the initial clustering center set: In the two-dimensional space after dimensionality reduction, select K points as the combination of initial clustering centers. This set is the benchmark for privacy risk assessment and is used for subsequent distance calculation and classification of data points;

[0019] 2) Calculate the minimum distance from the clustering center set: For each data point, calculate the distance from the entire clustering center set. Based on the distance from this set, classify the data points into areas close to or far from this set;

[0020] 3) Update the data distribution: Based on the distance between the data points and the clustering center set, redistribute the data points to ensure that the points close to the clustering center set are grouped into one category, and the points far from the clustering center set are grouped into another category;

[0021] 4) Iterative update: According to the calculation results, iteratively update the distribution and classification of the data points until the distribution of the data points is stable.

[0022] 5) Generate the clustering result: Finally, generate the privacy risk clustering result to provide risk classification and assessment information.

[0023] Step Three: Self-learning weight assignment mechanism

[0024] 1) Calculate the mean of the feature matrix: Based on the feature matrix of the selected K points, calculate its mean matrix to obtain a matrix reflecting the overall features, which is used to accurately evaluate the contribution of each risk factor to the overall privacy risk.

[0025] 2) Calculate the similarity: Based on the original feature matrix and the mean matrix, successively remove each row of the feature matrix, and recalculate the similarity between the two matrices after removal through the Mahalanobis distance.

[0026] 3) Mahalanobis distance calculation and weight assignment: Based on the calculation of the Mahalanobis distance, determine the contribution degree of each risk factor to the overall matrix. If the Mahalanobis distance increases after removing a certain row, it indicates that this row makes a greater contribution to the whole, and its weight should be adjusted smaller accordingly; conversely, if the Mahalanobis distance decreases, the weight of this row should be increased accordingly.

[0027] 4) Weight coefficient normalization: Finally, normalize all the calculated weight coefficients and convert them into values between 0 and 1 to form a 1×4 weight vector for subsequent calculations. Risk feature matrix and weight vector calculation: Calculate the generated 5×4 privacy risk feature matrix and the one-dimensional weight vector to obtain a specific privacy risk assessment score for evaluating the overall privacy risk level.

[0028] Advantages of the present invention: The present invention can more comprehensively and accurately evaluate the privacy risks of Android applications through multi-dimensional analysis and self-learning weight assignment mechanism, help developers effectively identify and prevent potential privacy leaks, and improve the overall data security. Brief Description of the Drawings

[0029] Att Figure 1 : The overall framework diagram of the privacy risk measurement method, describing the overall process from data preprocessing, risk matrix generation to clustering analysis and weight assignment.

[0030] Att Figure 2 : The abstract drawing of the privacy risk measurement method, showing the overall process from data collection and preparation, risk feature extraction, risk quantitative analysis to risk assessment report generation. Detailed Embodiments

[0031] Embodiment 1: Measurement of excessive permission requests

[0032] Extraction process:

[0033] 1. Use a static analysis tool to extract the permission declarations of Android applications.

[0034] 2. Analyze the matching situation between the extracted permissions and the application scenario requirements to determine whether the permission application exceeds the minimum necessary permissions.

[0035] Risk level classification:

[0036] Low risk: Only request basic and necessary permissions. Count the instances in the application that only request basic necessary permissions.

[0037] Medium risk: Request some non-necessary permissions. Count the instances in the application that request some non-necessary permissions.

[0038] High risk: A large number of unnecessary sensitive permissions are requested. Count the instances in the application where a large number of unnecessary sensitive permissions are requested.

[0039] Severe risk: Excessive sensitive permissions are requested, including reading and modifying sensitive data. Count the instances in the application where excessive sensitive permissions are requested.

[0040] Example 2: Verification of data sources

[0041] Analysis process:

[0042] 1. Use static analysis tools to trace data source nodes, including API calls or external inputs.

[0043] 2. Confirm the credibility and security of the data source, and evaluate whether the received data may lead to privacy leakage.

[0044] Risk level classification:

[0045] Low risk: The data comes from a trusted internal API or user input. Count the instances where the data comes from a trusted internal API or user input.

[0046] Medium risk: The data comes from partially trusted external APIs. Count the instances where the data comes from partially trusted external APIs.

[0047] High risk: The data comes from unknown or untrusted external APIs. Count the instances where the data comes from unknown or untrusted external APIs.

[0048] Severe risk: The data comes from clearly untrusted or malicious external APIs. Count the instances where the data comes from clearly untrusted or malicious external APIs.

[0049] Example 3: Analysis of data flow to the terminal

[0050] Tracking process:

[0051] 1. Determine the data flow, including internal storage or external servers.

[0052] 2. Check the security of the terminal to ensure that the data is not transmitted to high-risk or untrusted terminals.

[0053] Risk level classification:

[0054] Low risk: The data is stored locally or on a trusted internal server. Count the instances where the data is stored locally or on a trusted internal server.

[0055] Medium risk: The data flows to a trusted external server. Count the instances where the data flows to a trusted external server.

[0056] High risk: Data flows to multiple external servers, and some of these servers have low credibility. Count the instances where data flows to multiple external servers and some of the servers have low credibility.

[0057] Severe risk: Data flows to unknown or untrusted external servers. Count the instances where data flows to unknown or untrusted external servers.

[0058] Example 4: Detection of HTTP / HTTPS Channel Privacy Leakage and Verification of Secure Transmission

[0059] Detection process:

[0060] 1. Use static analysis tools to monitor the data transmission in the HTTP / HTTPS channel.

[0061] 2. Verify the encryption strength of the data transmission and detect whether sensitive information is transmitted through an insecure channel.

[0062] Risk level classification:

[0063] Low risk: Use industry-standard encryption methods and do not expose sensitive data. Standard HTTPS communication is used for non-sensitive information.

[0064] Medium risk: Use an encrypted channel, but there are minor vulnerabilities or outdated encryption methods are used. Use an old protocol or encrypted communication with small defects.

[0065] High risk: There are significant weaknesses in the encrypted channel, such as improper encryption algorithms or certificate verification. Due to encryption weaknesses, the communication is easily intercepted or decrypted.

[0066] Severe risk: The encrypted channel is highly vulnerable, or critical sensitive data is exposed due to encryption or implementation defects. Sensitive data is not encrypted or there are vulnerabilities, resulting in unauthorized access by attackers.

[0067] Example 5: Security Analysis of API Communication

[0068] Testing process:

[0069] 1. Use security testing tools to analyze the security of API communication.

[0070] 2. Detect whether appropriate encryption measures are used during the data transmission process to ensure that the data during the communication process will not be intercepted or tampered with.

[0071] Risk level classification:

[0072] Low risk: All API communications use strong encryption. Count the instances where all API communications are fully encrypted.

[0073] Medium risk: Most API communications are encrypted, and a small number are not. Count the instances where a small number of API communications are not encrypted.

[0074] High risk: Many API communications are not encrypted, and the amount of sensitive data involved is small. Count the instances where API communications are not encrypted and involve sensitive data.

[0075] Severe risk: A large amount of sensitive data is transmitted through unencrypted API communications. Count the instances where API communications are not encrypted and involve a large amount of sensitive data.

[0076] The specific process of risk feature matrix generation and K-means clustering analysis

[0077] Risk feature matrix generation:

[0078] 1. Data preprocessing: Assume there are n Android applications, and 5 key risk factors are extracted from each application, namely: excessive permission requests, data sources, data flow to the terminal, HTTP / HTTPS channel security, and API communication security. Each risk factor is extracted through a static analysis tool and processed into a numerical vector.

[0079] For each application i, its risk factors can be represented as a one-dimensional vector V i ∈R 5 , that is:

[0080] V i =[v i1 ,v i2 ,v i3 ,v i4 ,v i5

[0081] 2. Generate the risk feature matrix: The risk privacy vectors V i of each application i are combined into a matrix F i , as follows:

[0082]

[0083] 3. Matrix dimensionality reduction: Use principal component analysis (PCA) to perform dimensionality reduction on the risk matrix FFF, reducing the 5-dimensional feature matrix to a two-dimensional space for subsequent clustering analysis.

[0084] The matrix FPCA after dimensionality reduction is:

[0085] F PCA =PCA(F)

[0086] Here, the PCA dimensionality reduction converts the original feature matrix into a two-dimensional matrix by selecting the two principal components with the largest eigenvalues.

[0087] ​K-means clustering analysis:

[0088] 1. Select the initial cluster center set: In the two-dimensional space after dimensionality reduction, select K points as the initial cluster center set μ1, μ2, …, μ K ∈R 2 . These cluster centers serve as the benchmarks for privacy risk assessment.

[0089] 2. Calculate the minimum distance to the cluster center set: For each data point xi ∈ F PCA , calculate its distance to the entire cluster center set. Here, instead of calculating the distance to a single cluster center only, calculate the minimum distance to the entire cluster center set. Let the cluster center set be Centers, then for each data point xi, calculate its minimum distance to all points in the set:

[0090]

[0091] 3. Update the data distribution: Redistribute the data points based on the Mahalanobis distance between the data points and the cluster center set.

[0092]

[0093] Decision rule:

[0094]

[0095] Ensure that the points close to the cluster center set are grouped into one class, and the points far from the cluster center set are grouped into another class.

[0096] 4. Generate the clustering results: Finally, generate the clustering results C1, C2, …, C K , and each data point x i is assigned to a cluster. Through these clusters, generate the privacy risk classification and evaluation information of the application. Among them, C K represents the set of data points in the k-th cluster, which is used for privacy risk assessment.

[0097] Through these specific implementation steps and sufficient necessity explanations, a risk feature matrix can be constructed, and through K-means clustering analysis, a detailed and accurate assessment of the privacy protection level of mobile applications can be provided, helping developers and users make more informed decisions and enhancing the security of mobile applications and the trust of users.

Claims

1. A privacy risk measurement method for Android mobile applications, characterized in that, The method includes the following steps: Step (1) Factor extraction: Extract privacy risk factors in the Android application through a static analysis tool, including five key privacy risk factors: excessive permission requests, data sources, data flow to the terminal, detection of privacy leakage and verification of secure transmission in the HTTP / HTTPS channel, and API communication security analysis; Step (2) Generate a risk feature matrix: Convert the five extracted key privacy risk factors into numerical vectors and construct a privacy risk feature matrix; Step (3) Matrix dimensionality reduction: Perform dimensionality reduction processing on the generated privacy risk feature matrix through principal component analysis (PCA) to obtain its representation in a two-dimensional space; Step (4) K-means clustering analysis: In the reduced two-dimensional space, select K points as the initial clustering centers, redistribute the data points by calculating the Mahalanobis distance between each data point and the set of clustering centers, and finally generate a privacy risk clustering result; Step (5) Self-learning weight assignment: Automatically adjust the weight of each privacy risk factor according to the mean calculation of the feature matrix and the change in the Mahalanobis distance, generate a normalized weight vector, and calculate it with the risk feature matrix to obtain the final privacy risk assessment score.

2. The method according to claim 1, characterized in that The excessive permission request detection includes the following steps: Step (1) Use a static analysis tool to extract the permission declarations of the application; Step (2) Analyze whether the extracted permissions exceed the application scenario requirements and evaluate whether there are excessive requests in the permission application.

3. The method according to claim 1, characterized in that, The data source verification includes the following steps: Step (1) Trace the data source nodes through a static analysis tool, including API calls or external inputs; Step (2) Confirm the credibility and security of the data source and evaluate whether the data source may pose a privacy leakage risk.

4. The method according to claim 1, wherein The data flow to the terminal analysis includes the following steps: Step (1) Trace the path of the data flow to the terminal, including internal storage or external servers; Step (2) Evaluate the security of the terminal to ensure that the data will not be transmitted to high-risk or untrusted terminals.

5. The method according to claim 1, characterized in that, The detection of privacy leakage and verification of secure transmission in the HTTP / HTTPS channel includes the following steps: Step (1) Monitor the data transmission in the HTTP / HTTPS channel and detect whether sensitive information is transmitted through an insecure channel; Step (2) Verify whether industry-standard encryption methods are used during the transmission process and evaluate whether there are potential privacy leakage risks.

6. The method according to claim 1, wherein The API communication security analysis includes the following steps: Step (1) Analyze the security of API communication through a security testing tool; Step (2) Detect whether appropriate encryption measures are used during the data transmission process to prevent the data from being intercepted or tampered with.

7. The method according to claim 1, characterized in that The self-learning weight assignment includes the following steps: Step (1) Calculate the mean matrix based on the feature matrix of the selected K points; Step (2) Remove each row in the feature matrix in turn, calculate its contribution to the overall privacy risk through the Euclidean distance, and automatically adjust the weight of each privacy risk factor.

Citation Information

Patent Citations

  • An Android malicious software detection method based on a semi-supervised K-Means clustering algorithm

    CN109670310A

  • Gas pipeline corrosion stage prediction method based on K-means clustering-LSTM (Long Short Term Memory)

    CN116049707A