Android application third-party library privacy policy analysis and generation method
By combining static analysis and natural language processing technology for Android application third-party libraries, a third-party library privacy policy was generated, which solved the problem of missing generation of third-party library privacy policies, achieved more complete and accurate generation of privacy policies, and improved application compliance.
Patent Information
- Application Number
- CN202411952182.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the privacy policy generation of the third-party library is missing, resulting in the inability to fully disclose the personal information collected by the third-party library when the application privacy policy is compiled, which may lead to inconsistent with the specified privacy policy.
By static analysis of the third-party library of Android applications, the data flow of accessing user data is obtained, and combined with the data disclosure information in the application privacy policy collection, natural language processing technology is used to merge and deduplicate, and finally the privacy policy of the third-party library is output through the template.
The generated third-party library privacy policy is more complete and accurate, which can effectively solve the problem of incomplete collection of personal information by third-party library, and improve the compliance of application privacy policy.
Smart Images

Figure CN119989399A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of program static analysis and privacy protection, and in particular to a method for analyzing and generating a privacy policy for a third-party library of an Android application. Background Art
[0002] With the rapid development of mobile terminal devices, applications running on their operating systems have also brought great convenience to people's work and life. Among them, the Android operating system occupies more than 70% of the current global smartphone operating system market share, and applications developed and run on the Android operating system have also become an important part of Internet applications. At the same time, users are also increasingly concerned about the privacy issues of applications collecting personal information, so the compliance of standardized applications cannot be ignored. The third-party libraries used in the process of developing applications may lack privacy policies.
[0003] Application developers use many existing third-party libraries to simplify the development process when developing applications. If users want to know what personal information is collected by third-party libraries and the application itself, they can only view how user personal information is collected, used and disclosed in the application's privacy policy. When writing the application privacy policy, the author should disclose the personal information collected by the third-party libraries used. However, some third-party libraries do not provide privacy policies, and developers may not know the internal situation of the integrated third-party libraries, resulting in the author of the application privacy policy not knowing what personal information the third-party libraries will collect, which will make the written application privacy policy non-compliant.
[0004] At present, the work related to privacy policy mainly focuses on the privacy policy of the application. The privacy policy of the application can be directly obtained by analyzing the code of the application, and the design-related questions are answered by the developers to obtain the privacy-related information. At present, the privacy policy of the third-party library still has the following problems: there is no work on the generation of privacy policy for third-party libraries. First, there is no special platform to supervise the privacy policy provided by the third-party library, and it is difficult to collect the privacy policy of the third-party library; second, it is difficult to obtain the privacy access behavior of the third-party library. The third-party library is mainly provided in binary files, such as jar or aar files, and the functions of the third-party library cannot be executed independently; finally, many works related to third-party libraries are mainly obtained by analyzing the code of the application to obtain the path of personal information flowing to the third-party library, but the application may only use some functions in the third-party library, which leads to incomplete personal information collected by the third-party library. Summary of the invention
[0005] This application provides a method for analyzing and generating privacy policies for third-party libraries of Android applications, which can be used to solve the technical problem of incomplete personal information collected by third-party libraries.
[0006] This application provides a method for analyzing and generating a privacy policy for a third-party library of an Android application. The method takes an Android application privacy policy set and a sample application of a third-party library as input, and takes a generated third-party library privacy policy as an output result. The method includes:
[0007] Step 1: Take the sample application of the third-party library as input, then obtain the data stream of the third-party library accessing user data, and then obtain the data type and the third-party library based on the information in the data stream;
[0008] Step 2: extract data disclosure information about the third-party library from the application privacy policy set, merge and remove duplicates of the obtained user data disclosure information, and obtain the collected user data set of the third-party library;
[0009] Step 3, combining the user data results obtained in step 1 with the user data results obtained in step 2, and outputting the privacy policy of the third-party library through a template.
[0010] Furthermore, in step 1, the sample application of the third-party library is used as input to obtain the data stream of the third-party library accessing the user data, and then the data type and the third-party library are obtained according to the information in the data stream, including:
[0011] Step 1-1, taking the sample application of the third-party library as input, optimizing the data flow analysis by modifying the source method and the receiving method, and then obtaining a more comprehensive data flow of the third-party library accessing user data;
[0012] The data stream that passes the static taint analysis cannot be used directly because the source method and the receiving method may not be able to directly obtain the data type and entity, which can be divided into the following situations:
[0013] Source method: APIs access content providers, APIs request servers;
[0014] Receiving method: network-related APIs;
[0015] Step 1-2: Build a function call graph of the sample application, determine the data types and third-party libraries related to the data flow source method and receiving method through the call graph, and extract sensitive information in the data flow through reverse DNS and keyword similarity;
[0016] Solutions for situations where data types and entities cannot be obtained directly:
[0017] APIs access content providers: locate the API node in the call graph, start from the current node, traverse the call graph with reverse DNS, find the URIs that define the content provider and pass the URIs to all methods of the source method, and then use the mapping relationship in the ontology to map the URI to the data type;
[0018] APIs request to the server: locate the API node in the call graph, obtain sensitive strings from the parameters of the API call based on keyword similarity, and map them to data types in the ontology;
[0019] Network-related APIs: locate the API node in the call graph, record the string passed to the method as a parameter, if the string is in the form of a domain name, directly obtain the third-party entity through keyword extraction or the constructed mapping pair, if the string is an IP address without involving a domain name URL, perform a reverse DNS traversal to resolve the IP address into a domain name, and then determine the third-party entity; if there is no useful string, record the class name of the receiver method and match the class name with the third-party library;
[0020] Steps 1-3, create an ontology and mapping suitable for the third-party library to establish a complete third-party library data type, dynamically add new data type mappings as the data set increases, convert the obtained data type into the standardized data type in the ontology, and save the entity, collection, and data type as a user information triple to the third-party library to collect user data;
[0021] Construct the mapping relationship of the ontology. First, take the scope of user personal information given in the regulations as the data type in the initial ontology. Then expand the ontology by extracting user privacy information from the application privacy policy. Gradually construct the initial ontology, and continue to expand the ontology when encountering new data types in the future.
[0022] Further, as a specific example, as described in step 2, the data disclosure information about the third-party library is extracted from the application privacy policy set, and the obtained user data disclosure information is merged and deduplicated to obtain the collected user data set of the third-party library. The specific steps are as follows:
[0023] Step 2-1, first pre-process the application privacy policy set, remove the application privacy policies that do not meet the requirements, and then use natural language processing technology to extract the data disclosure information about the third-party library from the remaining application privacy policy set, and then merge them. Use natural language processing technology to extract new data types from the privacy policy, find the inclusive relationship between the data type and other data types in the ontology, and dynamically build the ontology edge and expand it into the ontology;
[0024] Step 2-2, remove duplicate privacy data in the merged information: Use semantic analysis methods to identify synonyms and context matching, identify the relationship between words, understand synonyms, near-synonyms, and context dependencies, and then deduplicate at the semantic level; third-party entities, personal information, permission information, purpose of use, and sharing method are taken as user information quintuples and saved in a third-party library to collect user data.
[0025] Furthermore, in step 3, the user data result obtained in step 1 is combined with the user data result obtained in step 2, and the third-party library privacy policy is output through a template, including:
[0026] Step 3-1, combining the user data results obtained in step 1 and the user data results obtained in step 2, and standardizing them into a five-tuple form after deduplication;
[0027] Step 3-2, formulate a policy template for the third-party library, input the data information in the form of quintuples into the template, and then output the complete privacy policy of the third-party library.
[0028] Compared with the prior art, the present invention has the following significant advantages: (1) In response to new problems, there is currently no work on generating privacy policies for third-party libraries. For the first time, it is proposed to deduplicate and merge third-party disclosure information in the application privacy policy set, and to perform static analysis on the sample applications of the third-party library, so that the generated third-party library collects more complete user data and the method is more effective; (2) It is a fully automatic static analysis plus natural language processing technology method, which does not require running the sample applications of the third-party library, and has the significant advantages of faster generation time and more complete user data; (3) This method can expand data types and ontology mapping to ensure accurate extraction of data disclosure information. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is an overview diagram of the privacy policy generation tool for the third-party library of Android applications provided by the present invention.
[0030] Figure 2 This is an example diagram of the static analysis data flow of an APK file.
[0031] Figure 3 This is an example diagram of a five-tuple of user privacy information collected by a third-party library. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0033] The following first introduces the embodiments of the present application in conjunction with the accompanying drawings.
[0034] This application provides a method for analyzing and generating privacy policies for third-party libraries of Android applications.
[0035] Taking the Android application privacy policy set and the sample application of the third-party library as input and the generated third-party library privacy policy as output, the method includes:
[0036] Step 1: Based on the PTPDroid tool, take the sample application of the third-party library as input, then obtain the data stream of the third-party library accessing user data, and then obtain the data type and third-party library according to the information in the data stream;
[0037] Step 2: extract data disclosure information about the third-party library from the application privacy policy set, merge and remove duplicates of the obtained user data disclosure information, and obtain the collected user data set of the third-party library;
[0038] Step 3, combining the user data results obtained in step 1 with the user data results obtained in step 2, and outputting the privacy policy of the third-party library through a template.
[0039] As a specific example, as described in step 1, based on the PTPDroid tool, the sample application of the third-party library is used as input to obtain the data stream of the third-party library accessing user data, and then the data type and the third-party library are obtained according to the information in the data stream. The specific steps are as follows:
[0040] Step 1-1, use the PTPDroid tool to take the sample application of the third-party library as input, optimize the data flow analysis by modifying the source method and the receiving method, and then obtain a more comprehensive data flow of the third-party library accessing user data;
[0041] The data stream that passes the static taint analysis cannot be used directly because the source method and the receiving method may not be able to directly obtain the data type and entity, which can be divided into the following situations:
[0042] Source method: APIs access content providers, APIs request data from servers
[0043] Receiving method: Network-related APIs
[0044] Step 1-2: Use the FlowDroid tool to build a function call graph of the sample application, determine the data types and third-party libraries related to the data flow source method and receiving method through the call graph, and extract sensitive information in the data flow through reverse DNS and keyword similarity;
[0045] Solutions for situations where data types and entities cannot be obtained directly:
[0046] APIs access content providers: locate the API node in the call graph, start from the current node, traverse the call graph with reverse DNS, find the URIs that define the content provider and pass the URIs to all methods of the source method, and then use the mapping relationship in the ontology to map the URI to the data type.
[0047] APIs request to the server: locate the API node in the call graph, obtain sensitive strings from the parameters of the API call based on keyword similarity, and map them to data types in the ontology.
[0048] Network-related APIs: Locate the API node in the call graph and record the string passed as a parameter to the method. If the string is in the form of a domain name, directly obtain the third-party entity through keyword extraction or constructed mapping pairs. If the string is an IP address without a domain name URL, perform reverse DNS traversal to resolve the IP address into a domain name, and then use the above method to determine the third-party entity; if there is no useful string, record the class name of the receiver method and match the class name with the third-party library.
[0049] Steps 1-3, create an ontology and mapping suitable for the third-party library to establish a complete third-party library data type. Dynamically add new data type mappings as the data set increases to improve the accuracy and completeness of recognition. Convert the obtained data type to the standardized data type in the ontology, and save the entity, collection, and data type as a user information triple to the third-party library to collect user data.
[0050] Construct the mapping relationship of the ontology. First, take the scope of user personal information given in the regulations as the data type in the initial ontology. Then expand the ontology by extracting user privacy information from the application privacy policy, gradually build the initial ontology, and continue to expand the ontology when encountering rare new data types.
[0051] As a specific example, as described in step 2, the data disclosure information about the third-party library is extracted from the application privacy policy set, and the obtained user data disclosure information is merged and deduplicated to obtain the collected user data set of the third-party library. The specific steps are as follows:
[0052] Step 2-1, first pre-process the application privacy policy set, remove the application privacy policies that do not meet the requirements, and then use natural language processing technology to extract the data disclosure information about the third-party library from the remaining application privacy policy set, and then merge them. Use natural language processing technology to extract new data types from the privacy policy, find the inclusive relationship between the data type and other data types in the ontology, and dynamically build the ontology edge and expand it into the ontology;
[0053] Step 2-2, remove duplicate privacy data in the merged information: Use semantic analysis methods to identify synonyms and match contexts, identify the relationship between words, understand synonyms, near-synonyms, context dependencies, etc., and then deduplicate from a semantic level to ensure more accurate deduplication results; third-party entities, personal information, permission information, purpose of use, and sharing methods are taken as a five-tuple of user information and saved in a third-party library to collect a collection of user data.
[0054] As a specific example, as described in step 3, the user data result obtained in step 1 and the user data result obtained in step 2 are combined, and the privacy policy of the third-party library is output through a template. The specific steps are as follows:
[0055] Step 3-1, combining the user data results obtained in step 1 and the user data results obtained in step 2, and standardizing them into a five-tuple form after deduplication;
[0056] Step 3-2, formulate a policy template for the third-party library, input the data information in the form of five-tuples into the template, ensure that the statement sentences comply with the regulations, and then output the complete privacy policy of the third-party library.
[0057] After running the tool, you can obtain the user data collected by the third-party library in the sample application of the third-party library and all non-duplicate disclosures of the user data collected by the third-party library in the application privacy policy set, and add the obtained user data collected by the third-party library into the provided template for output; if no user data is obtained in the result, directly output that the third-party library does not collect user data.
[0058] The present application is further described below in conjunction with specific embodiments.
[0059] The present invention is a method for automatically generating privacy policies for third-party libraries of Android applications. By analyzing the sample applications of the third-party libraries of Android applications and extracting the information disclosed by the third-party libraries in the application privacy policy set, the privacy policy results of the third-party libraries are generated. The specific work is as follows: Figure 1 shown.
[0060] First, identify a third-party library for which a privacy policy is to be generated, and perform static analysis on the sample application of the third-party library. Then, extract the disclosure information about the third-party library in the application privacy policy collection. Finally, combine the user information collected by the third-party library obtained earlier and add it to the template to obtain the privacy policy of the third-party library.
[0061] In this implementation, the method comprises the following steps:
[0062] Step 1: For a sample application of a third-party library of an Android application to be tested, identify the user privacy data access behavior of the third-party library through static analysis. The specific steps are as follows:
[0063] Step 1-1, use the PTPDroid tool to take the sample application of the third-party library as input, and then obtain the data flow of the third-party library accessing user privacy data, such as Figure 2 As shown, we get the data flow from getUserLocation() as the source method to sendLocationToAPI() as the receiving method;
[0064] Step 1-2, use the FlowDroid tool to build a function call graph for the sample application, and determine the data types and third-party libraries related to the data flow source method and receiving method through the call graph. As shown in Table 1, the data type obtained in the source method according to the pre-established mapping is location information, such as Figure 2 In the receiving method shown, it is found that the variable value https: / / www.tencent.com / is passed to the method, and based on the keyword mapping, it is determined that the entity is tencent.
[0065] Table 1
[0066]
[0067]
[0068] Table 1 is a mapping table between data types and class names and method names.
[0069] Step 1-3: save the entity, collection, and data type as a user information triplet to a third-party library to collect a collection of user privacy data.
[0070] Step 2: extract the privacy data disclosure information about the third-party library from the application privacy policy set, merge and remove duplicates of the obtained user data disclosure information, and obtain the collected user data set of the third-party library, as follows:
[0071] Step 2-1, applying the privacy policy set to extract the privacy data disclosure information about the third-party library using natural language processing technology, and extracting multiple pieces of privacy data disclosure information about the third-party library;
[0072] Step 2-2, merge the private data information and remove the duplicate private data in the merged information. Deduplication will normalize different forms of the same private data to avoid duplicate private data information. It will also standardize the Chinese and English permission information and save the third-party entity, personal information, permission information, purpose of use and sharing method as a five-tuple of user private information to the collection of user private data collected by the third-party library, such as Figure 3 shown.
[0073] Step 3: Combine the user privacy data results obtained in step 1 with the user privacy data results obtained in step 2, and output the privacy policy of the third-party library through the template, as follows:
[0074] Step 3-1, such as Figure 3 As shown, the user data result obtained in step 1 and the user data result obtained in step 2 are combined, and after deduplication, they are standardized into a five-tuple form;
[0075] Step 3-2, input the data information in the form of quintuple into the template, and then output the complete third-party library privacy policy.
[0076] In summary, the present invention can effectively and efficiently generate a privacy policy for an Android application third-party library.
[0077] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A method for analyzing and generating privacy policy of a third-party library of Android applications, characterized in that: The method takes the Android application privacy policy set and the sample application of the third-party library as input, and takes the generated third-party library privacy policy as output. The method includes: Step 1: Take the sample application of the third-party library as input, then obtain the data stream of the third-party library accessing user data, and then obtain the data type and the third-party library based on the information in the data stream; Step 2: extract data disclosure information about the third-party library from the application privacy policy set, merge and remove duplicates of the obtained user data disclosure information, and obtain the collected user data set of the third-party library; Step 3, combining the user data results obtained in step 1 with the user data results obtained in step 2, and outputting the privacy policy of the third-party library through a template.
2. The method according to claim 1, characterized in that In step 1, the sample application of the third-party library is used as input to obtain the data stream of the third-party library accessing user data, and then the data type and the third-party library are obtained based on the information in the data stream, including: Step 1-1, taking the sample application of the third-party library as input, optimizing the data flow analysis by modifying the source method and the receiving method, and then obtaining a more comprehensive data flow of the third-party library accessing user data; The data stream that passes the static taint analysis cannot be used directly because the source method and the receiving method may not be able to directly obtain the data type and entity, which can be divided into the following situations: Source method: APIs access content providers, APIs request servers; Receiving method: network-related APIs; Step 1-2: Build a function call graph of the sample application, determine the data types and third-party libraries related to the data flow source method and receiving method through the call graph, and extract sensitive information in the data flow through reverse DNS and keyword similarity; Solutions for situations where data types and entities cannot be obtained directly: APIs access content providers: locate the API node in the call graph, start from the current node, traverse the call graph with reverse DNS, find the URIs that define the content provider and pass the URIs to all methods of the source method, and then use the mapping relationship in the ontology to map the URI to the data type; APIs request to the server: locate the API node in the call graph, obtain sensitive strings from the parameters of the API call based on keyword similarity, and map them to data types in the ontology; Network-related APIs: locate the API node in the call graph and record the string passed to the method as a parameter. If the string is in the form of a domain name, directly obtain the third-party entity through keyword extraction or the constructed mapping pair. If the string is an IP address without a domain name URL, perform a reverse DNS traversal to resolve the IP address into a domain name, and then determine the third-party entity. If there is no useful string, record the class name of the receiver method and match the class name with the third-party library. Steps 1-3, create an ontology and mapping suitable for the third-party library to establish a complete third-party library data type, dynamically add new data type mappings as the data set increases, convert the obtained data type into the standardized data type in the ontology, and save the entity, collection, and data type as a user information triple to the third-party library to collect user data; Construct the mapping relationship of the ontology. First, take the scope of user personal information given in the regulations as the data type in the initial ontology. Then expand the ontology by extracting user privacy information from the application privacy policy. Gradually construct the initial ontology, and continue to expand the ontology when encountering new data types in the future.
3. The method according to claim 1, characterized in that As a specific example, as described in step 2, the data disclosure information about the third-party library is extracted from the application privacy policy set, and the obtained user data disclosure information is merged and deduplicated to obtain the collected user data set of the third-party library. The specific steps are as follows: Step 2-1, first pre-process the application privacy policy set, remove the application privacy policies that do not meet the requirements, and then use natural language processing technology to extract the data disclosure information about the third-party library from the remaining application privacy policy set, and then merge them. Use natural language processing technology to extract new data types from the privacy policy, find the inclusive relationship between the data type and other data types in the ontology, and dynamically build the ontology edge and expand it into the ontology; Step 2-2, remove duplicate private data in the merged information: Use semantic analysis methods to identify synonyms and context matching, identify the relationship between words, understand synonyms, near synonyms, and context dependencies, and then remove duplicates at the semantic level; The third-party entity, personal information, permission information, purpose of use and sharing method are taken as a five-tuple of user information and saved in a third-party library to collect a collection of user data.
4. The method according to claim 1, characterized in that: In step 3, the user data results obtained in step 1 are combined with the user data results obtained in step 2, and the privacy policy of the third-party library is output through a template, including: Step 3-1, combining the user data results obtained in step 1 and the user data results obtained in step 2, and standardizing them into a five-tuple form after deduplication; Step 3-2, formulate a policy template for the third-party library, input the data information in the form of quintuples into the template, and then output the complete privacy policy of the third-party library.