Method and apparatus for checking data acquisition disclosure in application, device, and product

By using deep learning language models to identify runtime page information in mobile applications, the problem of detecting explicit data acquisition during runtime has been solved, achieving more accurate and comprehensive detection of explicit data acquisition and ensuring the effective protection of users' right to know.

WO2026156842A1PCT designated stage Publication Date: 2026-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-01-26
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing mobile applications struggle to accurately detect explicit data acquisition information provided during user interactions, resulting in a failure to effectively protect users' right to know. In particular, detection schemes for explicit data acquisition information during runtime suffer from false positives and false negatives.

Method used

Using a deep learning-based language model, runtime page information is obtained by simulating user interaction, generating explicit recognition results, identifying whether the page information includes explicit data acquisition, and determining whether the data acquisition behavior meets the predetermined requirements based on the target data and the explicit recognition results.

Benefits of technology

It improves the accuracy and comprehensiveness of data acquisition disclosure detection, ensures the effective protection of users' right to know, reduces false alarms and false negatives, and enhances the ability to detect explicit data acquisition during runtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075319_30072026_PF_FP_ABST
    Figure CN2025075319_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for checking information acquisition disclosure, a device, and a product. The method comprises determining that an application acquires target data. The method further comprises acquiring page information on a runtime page in the application. The method further comprises using, on the basis of the page information, a language model to generate a disclosure identification result, the disclosure identification result indicating whether the page information comprises data acquisition disclosure. In addition, the method further comprises determining, on the basis of the target data and the disclosure identification result, whether acquisition of the target data meets a predetermined requirement. In this way, the accuracy of data acquisition disclosure identification can be improved, and the comprehensiveness of data acquisition disclosure checking can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, equipment and products for detecting explicit data acquisition in applications. Technical Field

[0001] This disclosure relates to the field of data security, and more specifically to methods, apparatus, devices and products for detecting explicit data acquisition in applications. Background Technology

[0002] As people become increasingly aware of data protection, users are paying more and more attention to how applications collect and process specific categories of target data. Therefore, there are regulations and requirements governing the acquisition of such target data, which clearly emphasize the importance of protecting users' right to know. Mobile applications, as technological carriers widely integrated into all aspects of people's lives and work, carry a large amount of target data, thus becoming a crucial and indispensable element in the governance of user right to know protection.

[0003] Explicit data acquisition refers to clearly informing users of the target data to be acquired and the purpose of acquiring that data within the application, and obtaining the user's consent. This is a transparent data processing mechanism designed to ensure that users have a full understanding of how the target data is acquired and used. Summary of the Invention

[0004] In a first aspect of the embodiments of this disclosure, a method for detecting explicit data acquisition in an application is provided. The method includes determining that the application acquires target data. The method also includes acquiring page information on a runtime page within the application. The method further includes generating an explicit acquisition result based on the page information using a language model, the explicit acquisition result indicating whether the page information includes explicit data acquisition. Furthermore, the method includes determining whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit acquisition result.

[0005] In a second aspect of the embodiments of this disclosure, an apparatus for detecting explicit data acquisition in an application is provided. The apparatus includes an acquisition behavior determination module configured to determine that the application acquires target data. The apparatus also includes a page information acquisition module configured to acquire page information on a runtime page in the application. Furthermore, the apparatus includes an explicit acquisition recognition module configured to generate an explicit acquisition result based on the page information using a language model, the explicit acquisition result indicating whether the page information includes explicit data acquisition. Additionally, the apparatus includes an explicit acquisition detection module configured to determine whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit acquisition result.

[0006] In a third aspect of embodiments of this disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement a method for detecting explicit data acquisition in an application. The method includes determining that the application acquires target data. The method also includes acquiring page information on a runtime page within the application. The method further includes generating an explicit identification result based on the page information using a language model, the explicit identification result indicating whether the page information includes explicit data acquisition. Furthermore, the method includes determining whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit identification result.

[0007] In a fourth aspect of embodiments of this disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to implement a method for detecting explicit data acquisition in an application. The method includes determining that the application acquires target data. The method also includes acquiring page information on a runtime page within the application. The method further includes generating an explicit identification result based on the page information using a language model, the explicit identification result indicating whether the page information includes explicit data acquisition. Furthermore, the method includes determining whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit identification result.

[0008] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 illustrates a schematic diagram of an example environment in which various embodiments of the present disclosure may be implemented;

[0011] Figure 2 shows a flowchart of a method for obtaining explicit data in a detection application according to some embodiments of the present disclosure;

[0012] Figure 3 illustrates a schematic diagram of the architecture of an example system for data acquisition in a detection application according to some embodiments of the present disclosure;

[0013] Figure 4 shows a schematic diagram of an example of generating prompt words that cause a language model to output explicit recognition results, according to some embodiments of the present disclosure;

[0014] Figure 5 shows a schematic diagram of an example of generating prompt words that cause a language model to output explicit components, according to some embodiments of the present disclosure;

[0015] Figure 6 shows a schematic diagram of an example of a language model for training explicit recognition results according to some embodiments of the present disclosure;

[0016] Figure 7 illustrates a schematic diagram of an example of training a language model for outputting explicit components according to some embodiments of the present disclosure;

[0017] Figure 8 shows a block diagram of an apparatus for data acquisition in a detection application according to some embodiments of the present disclosure; and

[0018] Figure 9 shows a block diagram of a device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0019] It is understood that all user-related data involved in this technical solution should be obtained and used only after authorization from the user. This means that if it is necessary to use a user's personal information in this technical solution, the user's explicit consent and authorization are required before obtaining this data; otherwise, no related data collection and use will be carried out. It should also be understood that when implementing this technical solution, relevant laws and regulations should be strictly followed in the process of data collection, use, and storage, and necessary technical measures should be taken to protect user data security and ensure the secure use of data.

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0022] Currently, mobile applications primarily rely on separate data acquisition statement documents (e.g., privacy policy documents) to inform users about the scope, methods, and purposes of their collection and processing of target data. Therefore, existing detection schemes for ensuring users' right to know focus on whether the mobile application's statement document complies with predetermined requirements.

[0023] However, users often ignore or skim through data acquisition statements, meaning they don't actually see the statements related to acquiring the target data. Compared to data acquisition statements, data acquisition prompts provided by mobile applications during actual user interaction ensure that users see the relevant content, thus more effectively protecting their right to know. This article refers to this type of data acquisition statement displayed during actual user interaction with the application as "runtime data acquisition explicitness." Currently, relevant data acquisition explicitness detection technologies typically neglect the detection of runtime data acquisition explicitness.

[0024] Furthermore, data acquisition declaration documents are easily accessible (e.g., readily available by accessing a link to the declaration document for a specific mobile application in an app store) and have a complete sentence structure, making them easy to parse. In contrast, runtime data acquisition declarations within mobile applications are distributed throughout various interactions between the mobile application and the user, making them difficult to obtain. Moreover, runtime data acquisition declarations are typically composed of scattered words or phrases. This fragmented structure makes runtime data acquisition declarations difficult to parse accurately, rendering relevant detection schemes unsuitable for direct detection of runtime data acquisition declarations within mobile applications.

[0025] Therefore, embodiments of this disclosure provide a scheme for detecting explicit data acquisition in an application. In this scheme, a computing device can determine that an application has performed an action to acquire target data. Furthermore, the computing device can acquire page information from runtime pages within the application. Then, based on the page information, the computing device can use a language model to generate an explicit acquisition result, indicating whether the page information includes explicit data acquisition. The computing device can then determine whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit acquisition result.

[0026] In this way, the language model can identify whether the page information displayed on the runtime page is explicit data acquisition, reducing false and false identifications and thus improving the accuracy of explicit data acquisition identification. Furthermore, by detecting data acquisition behavior in the application, it can determine whether the data acquisition behavior meets predetermined requirements based on the explicit identification results of the runtime page, thereby improving the comprehensiveness of explicit data acquisition detection.

[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. As shown in Figure 1, environment 100 includes a computing device 102, which can be any device with computing or processing capabilities. For example, computing device 102 can be a local server, cloud server, desktop computer, laptop computer, tablet computer, etc. In environment 100, application 104 is a running application, which can run on computing device 102 or on a device other than computing device 102 (e.g., a user device). Application 104 can be, for example, a content distribution application, a social application, a map application, a lifestyle service application, etc.

[0028] Computing device 102 can simulate user interaction with application 104 and acquire data associated with application 104 during the interaction (e.g., screenshots, data transmitted over a network, etc.). In environment 100, computing device 102 can determine that application 104 is acquiring target data 106 during its operation by monitoring data associated with application 104. Target data 106 is a specific type of data, and application 104 must meet predetermined requirements when acquiring this specific type of data. These predetermined requirements may specify that application 104 can acquire and use this specific type of data with the user's knowledge or consent. In some embodiments, computing device 102 can acquire system interfaces called by application 104 and determine whether the system interface is used to acquire target data 106. In some embodiments, computing device 102 can acquire data transmitted by application 104 over a network and determine whether the data contains target data 106.

[0029] In environment 100, computing device 102 can access runtime page 108 displayed in application 104. Runtime page 108 can be any page displayed by application 104 during operation, including pages that explicitly indicate data access. In environment 100, runtime page 108 can include page information 110, and an example of page information 110 is shown in Figure 1. In this example, page information 110 includes text content 112, button 114, and button 116. Text content 112 asks the user whether application A is allowed to access target data (e.g., target data 106) and informs the user of the purpose of accessing the target data (e.g., to provide precise services). Buttons 114 and 116 are also part of page information 110, providing the user with the option to allow application A to access the target data or deny application A access to the target data. Computing device 102 can utilize technologies such as optical character recognition (OCR) to extract page information 110 from runtime page 108.

[0030] In environment 100, after acquiring page information 110, the computing device can generate an explicit recognition result 120 based on the page information 110 using a language model 118. The explicit recognition result 120 can indicate whether the page information 110 includes explicit data acquisition. The language model 118 can be a deep learning-based natural language processing model that learns from massive amounts of text data, capturing the syntax, semantics, and contextual relationships of language, enabling it to understand natural language and generate high-quality, coherent text. The core working principle of the language model 118 is based on a Transformer architecture, using a self-attention mechanism to analyze the relationships between words and sentences in the input text to identify deep-level structures in the language. For example, the language model 118 can be a Large Language Model (LLM) such as GPT or BERT.

[0031] In some embodiments, computing device 102 can generate a prompt for language model 118 based on page information 110. This prompt indicates whether the page information contains explicit data acquisition information based on the provided page information. Computing device 102 can then generate an explicit recognition result 120 by inputting the generated prompt into language model 118. For example, explicit recognition result 120 can be "yes" or "no," where "yes" indicates that page information 110 contains explicit data acquisition information, and "no" indicates that page information 110 does not contain explicit data acquisition information. Due to the advantages of language model 118 in semantic and contextual understanding, even if the content of page information 110 consists of scattered words or phrases, language model 118 can still accurately understand the meaning of page information 110 and output explicit recognition result 120, reducing misidentification and missed identification.

[0032] In environment 100, computing device 102 can determine whether the acquisition behavior of target data 106 meets predetermined requirements based on target data 106 and explicit identification result 120. As shown in FIG1, explicit detector 122 can determine detection result 124 based on target data 106 and explicit identification result 120. Detection result 124 can indicate whether the acquisition behavior of application 104 of target data 106 meets predetermined requirements. For example, when computing device 102 detects that application 104 acquires target data 106 during operation, but does not detect a runtime page with explicit identification result 120 as "yes", it can determine that application 104 did not provide the user with corresponding data acquisition explicitness when acquiring target data 106, and therefore does not meet predetermined requirements. If computing device 102 detects a runtime page with explicit identification result 120 as "yes", it can further determine whether the data acquisition explicitness in page information 110 is associated with target data 106, and whether the data acquisition explicitness meets predetermined requirements for target data 106.

[0033] In this way, language model 118 can identify whether the page information 110 displayed on runtime page 108 is explicit data acquisition, reducing false and false identifications and thus improving the accuracy of explicit data acquisition identification. Furthermore, by detecting data acquisition behavior in application 104, it can determine whether the data acquisition behavior meets predetermined requirements based on the explicit identification result 120 of runtime page 108, thereby improving the comprehensiveness of explicit data acquisition detection.

[0034] Figure 2 illustrates a flowchart of a method 200 for detecting explicit data acquisition in an application according to some embodiments of the present disclosure. Method 200 can be performed by a computing device, such as computing device 102 in Figure 1. As shown in Figure 2, at block 202, the computing device can determine that an application is acquiring target data. For example, in environment 100 as shown in Figure 1, computing device 102 can simulate a user interacting with application 104 and acquiring data associated with application 104 during the interaction. Computing device 102 can determine that application 104 is acquiring target data 106 during operation by monitoring the data associated with application 104. Target data 106 is a specific type of data, and application 104 must meet predetermined requirements when acquiring this specific type of data. These predetermined requirements may specify that application 104 can acquire and use this specific type of data with the user's knowledge or consent.

[0035] In box 204, the computing device can acquire page information from a runtime page within an application. For example, in environment 100 as shown in FIG1, computing device 102 can acquire a runtime page 108 displayed in application 104. Runtime page 108 can be any page displayed by application 104 during runtime, including pages explicitly used to display data acquisition. Computing device 102 can utilize technologies such as OCR to extract page information 110 from runtime page 108.

[0036] In box 206, the computing device can generate an explicit recognition result based on page information using a language model. This explicit recognition result indicates whether the page information includes explicit data acquisition information. For example, in environment 100 as shown in FIG1, the computing device can generate an explicit recognition result 120 based on page information 110 using a language model 118. This explicit recognition result 120 indicates whether the page information 110 includes explicit data acquisition information. In some embodiments, the computing device 102 can generate a prompt word for the language model 118 based on page information 110. This prompt word indicates whether the page information contains explicit data acquisition information based on the provided page information. The computing device 102 can then generate the explicit recognition result 120 by inputting the generated prompt word into the language model 118. For example, the explicit recognition result 120 can be "yes" or "no," where "yes" indicates that page information 110 contains explicit data acquisition information, and "no" indicates that page information 110 does not contain explicit data acquisition information.

[0037] In box 208, the computing device can determine whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit identification result. For example, in environment 100 as shown in FIG1, computing device 102 can determine whether the acquisition behavior of target data 106 meets predetermined requirements based on target data 106 and explicit identification result 120. For example, when computing device 102 detects that application 104 acquires target data 106 during operation, but does not detect a runtime page with explicit identification result 120 as "yes", it can determine that application 104 did not provide the user with corresponding data acquisition explicitness when acquiring target data 106, and therefore does not meet predetermined requirements. If computing device 102 detects a runtime page with explicit identification result 120 as "yes", it can further determine whether the data acquisition explicitness in page information 110 is associated with target data 106, and whether the data acquisition explicitness meets predetermined requirements for target data 106.

[0038] In this way, the language model can identify whether the page information displayed on the runtime page is explicit data acquisition, reducing false and false identifications and thus improving the accuracy of explicit data acquisition identification. Furthermore, by detecting data acquisition behavior in the application, it can determine whether the data acquisition behavior meets predetermined requirements based on the explicit identification results of the runtime page, thereby improving the comprehensiveness of explicit data acquisition detection.

[0039] In some embodiments, when determining that an application is acquiring target data, the computing device may acquire a system interface invoked by the application, and then determine that the application is acquiring the target data by determining that the system interface is used to acquire the target data. Alternatively, the computing device may acquire network data sent by the application, and then determine that the application is acquiring the target data based on determining that the network data includes the target data.

[0040] In some embodiments, when obtaining page information from the runtime page of the application, the computing device can log in to the application by logging into a third-party account and using single sign-on. Then, the computing device can obtain the page information from the runtime page of the application while logged in.

[0041] In some embodiments, when determining whether the acquisition of target data meets predetermined requirements, in response to the explicit identification result indicating that the page information includes explicit data acquisition information, the computing device can determine that the page information is associated with the target data. Then, the computing device can determine the acquisition time of the target data and the display time of the runtime page. In response to the acquisition time of the target data being earlier than the display time of the runtime page, the computing device can determine that the acquisition of the target data does not meet predetermined requirements.

[0042] In some embodiments, when determining that target data is associated with page information, the computing device may determine a first data category of the target data and a second data category associated with the page information. The computing device may then determine the semantic relevance between the first and second data categories, whereby the semantic relevance indicates whether the first and second data categories are semantically similar or semantically hierarchical. The computing device may then determine the association between the target data and the page information based on this semantic relevance. In some embodiments, in response to determining that page information is associated with target data, the computing device may, based on the page information, utilize a second language model to determine whether the acquisition of the target data meets predetermined requirements.

[0043] Figure 3 illustrates a schematic diagram of the architecture of an example system 300 for detecting explicit data acquisition in an application according to some embodiments of the present disclosure. As shown in Figure 3, system 300 includes a dynamic tester 304, a first language model 312, a second language model 318, and an explicit detector 328.

[0044] The Dynamic Tester 304 can be configured to use dynamic testing techniques combined with single sign-on to enable automatic application login. This fully triggers the business logic of the application under test in the logged-in state, expanding the scenarios that can trigger target data acquisition events and explicit data acquisition events, thereby improving the comprehensiveness of the test.

[0045] The target data acquisition identifier 320 can be configured to capture the target data acquisition behavior performed by the application by monitoring system interface calls associated with the target data and identifying data sent by the application over the network.

[0046] The first language model 312 and the second language model 318 can constitute a language model-based explicit parser. The first language model 312 can be configured to detect whether page information in a runtime page includes explicit data acquisition. The second language model 318 can be configured to extract explicit components (e.g., explicit components specified by a predetermined requirement) from the page information of the runtime page. In embodiments of this disclosure, pre-trained language models can be fine-tuned using collected offline data to generate the first language model 312 and the second language model 318. It should be noted that in some embodiments, the first language model 312 and the second language model 318 can be the same language model, which can simultaneously integrate the functionality of both the first language model 312 and the second language model 318.

[0047] The explicit detector 328 can be configured to construct a hypothetical ontology mapping between the target data acquisition behavior and the runtime data acquisition explicitness, thereby associating them. Then, the explicit detector 328 can determine, based on the association between the target data acquisition behavior and the runtime data acquisition explicitness, whether runtime data acquisition explicitness for the target data acquisition behavior is missing in the application under test. Furthermore, the explicit detector 328 can also determine, based on this association, whether the explicit components contained in the data acquisition explicitness meet the requirements for explicit components specified in the predetermined requirements.

[0048] As shown in Figure 3, the dynamic tester 304 can simulate user interaction with the application under test 302, enabling the system 300 to monitor various target data acquisition behaviors and runtime data acquisition indications within the application 302. In mobile applications, account systems are widely used to facilitate the management of user-related information and to provide services to users. This means that a large amount of interactive business logic in the application is only presented after the user logs into their account. Therefore, to trigger target data acquisition behaviors and present runtime pages indicating data acquisition more comprehensively and frequently, the dynamic tester 304 can automatically implement application login and subsequent dynamic testing.

[0049] Because different applications can maintain their own independent account systems, the same account information cannot be used to log in to different applications. However, to alleviate the inconvenience of managing multiple accounts for users and to facilitate integration with popular application platforms, most applications offer single sign-on (SSO), allowing users to log in to different applications with a single account. Therefore, the dynamic tester 304 can utilize SSO for dynamic testing. This approach improves the comprehensiveness of explicit data acquisition detection.

[0050] After logging into the application, the dynamic tester 304 can automatically simulate user interaction with the application and save page information during the interaction process. For example, the dynamic tester 304 can save a screenshot 306 of a runtime page in the application, and then use technologies such as OCR to extract page information 308 from the screenshot 306. In some embodiments, in addition to the screenshot 306, the dynamic tester 304 can also save the layout information of the runtime page and use the layout information to help extract page information 308 from the screenshot 306, thereby improving the completeness and semantic coherence of the page information 308.

[0051] Furthermore, during dynamic testing after logging into the application, the target data acquisition identifier 320 can continuously monitor target data acquisition behavior within the application. Upon detecting target data acquisition behavior, the identifier 320 can acquire the target data 322 corresponding to that behavior and the acquisition time 324 of the target data 322. When an application acquires target data, it typically calls system interfaces provided by the operating system. Additionally, even if some target data is not acquired through system interface calls, this data may be sent to the application's server via the network. Therefore, the identifier 320 can acquire system interface calls and network transmission data during application operation and determine whether target data acquisition behavior exists based on these data. This approach improves the comprehensiveness of target data acquisition behavior detection.

[0052] In some embodiments, the target data acquisition identifier 320 can acquire a pre-determined set of keywords representing target data. Then, the target data acquisition identifier 320 can determine words semantically similar to the keywords in the keyword set based on their hyponyms and sigmas, and add these words to the keyword set. For example, the target data acquisition identifier 320 can utilize the ConceptNet vocabulary database to determine words semantically similar to the keywords in the keyword set. The ConceptNet vocabulary database stores semantic relevance between words, determined based on semantic similarity and hyponyms between words. The target data acquisition identifier 320 can use the expanded keyword set to determine whether the network transmission data includes fields representing target data. Furthermore, the target data acquisition identifier 320 can also use regular expressions to identify data in the network transmission data that conforms to a specific regular expression (e.g., data conforming to an email address format) as target data, thereby determining the presence of target data acquisition behavior in the application. In this way, the comprehensiveness of target acquisition behavior detection can be improved.

[0053] As shown in Figure 3, after acquiring page information 308, system 300 can generate a first prompt word 310 based on page information 308. The first prompt word 310 is used to enable the first language model 312 to generate an explicit recognition result 314 based on page information 308. The explicit recognition result 314 indicates whether page information 308 includes explicit data acquisition. The first language model 312 can be, for example, the language model 118 in Figure 1. Because runtime data acquisition explicit statements have fragmented sentence structures and dispersed components, traditional schemes that identify explicit statements based on complete sentence structures and sentence semantics have low accuracy. However, since runtime data acquisition explicit statements are designed to attract user attention and allow users to promptly perceive and read the explicit content, the contextual semantics of runtime data acquisition explicit statements are usually relatively clear and fixed. Based on this, system 300 can use the overall text information of the runtime page as the underlying feature and utilize a language model to achieve semantic understanding of the page information, thereby determining whether the page information includes runtime data acquisition explicit statements. In this way, System 300 can reduce false alarms and false negatives caused by the traditional method of determining whether page information contains explicit data acquisition information by detecting the presence of specific explicit data acquisition components.

[0054] As shown in Figure 3, after generating the explicit identification result 314, if the explicit identification result 314 indicates that the page information 308 contains explicit data acquisition information, the system 300 can generate a second prompt word 316 based on the page information 308. The second prompt word 316 is used to enable the second language model 318 to extract explicit components 326 from the page information 308. For example, some predetermined requirements stipulate that explicit data acquisition information should include the following explicit components: the identity of the data controller initiating the target data processing behavior (i.e., IC), the rights enjoyed by the user (i.e., UR), the type of target data being processed (i.e., TD), the purpose of processing (i.e., PP), the basis for processing the target data (i.e., LB), the storage period of the target data being processed (i.e., SP), and the recipient of the target data being processed (i.e., ER). Therefore, the second prompt word 316 can require the second language model 318 to extract the above-mentioned explicit components from the provided page information and output the extracted results in the format of fields, with key-value pairs in the dictionary corresponding to the explicit components. In addition, if some explicit components are missing in the page information, the value of the explicit component can be set to empty. In this way, system 300 can determine whether the data acquisition declaration in page information 308 meets the predetermined requirements based on the explicit component 326.

[0055] In system 300, the explicit detection detector 328 can generate an explicit detection result 330 based on the target data 322, the acquisition time 324, and the explicit components 326. The explicit detection result 330 can indicate whether the data acquisition explicitness in the page information 308 meets predetermined requirements. In some embodiments, the explicit detection result 330 may include explicit content detection results, display timing detection results, and display format detection results. The explicit content detection result can indicate whether the data acquisition explicitness in the page information 308 includes all explicit components in the predetermined requirements (e.g., the seven explicit components mentioned above). If the explicit component 326 lacks some of the explicit components in the predetermined requirements, the explicit detection result 330 can indicate that the target data acquisition behavior does not meet the predetermined requirements or that the explicit content of the data acquisition explicitness does not meet the predetermined requirements.

[0056] The timing of the display can indicate whether explicit data acquisition is provided before or at the same time as the target data acquisition behavior. That is, whether the display time of page information 308 is earlier than or the same as the acquisition time of target data 322. If the display time of page information 308 is later than the acquisition time of target data 322, the explicit acquisition detection result 330 can indicate that the target data acquisition behavior does not meet the predetermined requirements or the timing of the data acquisition explicit display does not meet the predetermined requirements.

[0057] The display format detection result can indicate whether the explicit component 326 is clearly displayed on the runtime page. If the display of the explicit component 326 is unclear, the explicit detection result 330 can indicate that the target data acquisition behavior does not meet the predetermined requirements or that the display format of the data acquisition explicit does not meet the predetermined requirements.

[0058] Because the granularity of the target data acquisition behavior and the explicit data acquisition statement differs in representing or describing the target data, a direct association between the two is difficult. For example, when the collection of data A is detected, the corresponding explicit statement might be "collecting data B," where data B is a superordinate term of data A. This difference makes it difficult to directly establish a connection between the target data acquisition behavior and the explicit component of the data acquisition statement. However, despite the inconsistency in the granularity of their expressions, the target data acquisition behavior and the explicit data acquisition statement often have a semantic relationship of superordinate and hyperordinate terms. This actually reflects a conceptual inclusion relationship between the two different expressions, from a broader to a narrower extension.

[0059] Based on this, system 300 can use a semantic association method based on hyponym relationships to associate the target data acquisition behavior with its corresponding explicit data acquisition indication. After completing the association, the explicit indication detector 328 can determine that there is a target data acquisition behavior but no corresponding explicit data acquisition behavior, which does not meet the predetermined requirements. Furthermore, for cases where there is a corresponding explicit data acquisition indication, the explicit indication detector 328 can determine whether the explicit data acquisition indication itself is missing necessary explicit components based on the explicit component 326 generated by the second language model 318, and determine whether there are quality defects by analyzing the display time and content description of the explicit data acquisition indication.

[0060] In system 300, because hypernyms typically have broader or more ambiguous concepts, they appear relatively frequently. Based on this, in some embodiments, system 300 can utilize frequency analysis techniques to filter the expressions of all target data type components (i.e., TDs) collected in the explicit data acquisition training set, and determine the high-frequency expressions as the initial hypernym ontology. Subsequently, the filtered hypernym ontology can be associated with keywords in the target data keyword set. In this way, the explicit detector 328 can associate the target data acquisition behavior with the corresponding explicit data acquisition based on this association. This improves the accuracy of the association.

[0061] Regarding the display format, since ambiguous words appear more frequently, the explicit detector 328 can use frequency analysis to filter out the most common expressions of different explicit components. Then, by accepting user input, it can be determined whether the most common expressions are ambiguous, thus forming a fuzzy word set. The explicit detector 328 can then use this fuzzy word set for matching. If the expression of explicit component 326 matches a keyword in the fuzzy word set, it can be determined that the explicit component does not meet the predetermined requirements in terms of display format. For example, if the explicit component is the purpose of acquiring the target data (i.e., PP), an example of a fuzzy expression could be "for specific / market / analysis purposes." If the explicit component is the identity of the data controller (i.e., IC), an example of a fuzzy expression could be "partner" or "strategic partner." If the explicit component is the recipient of the target data (i.e., ER), an example of a fuzzy expression could be "third-party organization" or "partner." In this way, the accuracy of detecting the display format of explicit components can be improved.

[0062] In some embodiments, when generating a first prompt word for a first language model (e.g., first prompt word 310 in FIG3), the computing device may obtain a first task description, which indicates whether there is explicit data acquisition information in the provided page information that informs the user that target data has been acquired and that the target data is used for a specific purpose. Furthermore, the computing device may obtain a first example, including a page information example and an explicit recognition result example. The computing device may generate the first prompt word based on the page information, the first task description, and the first example. The computing device can then generate an explicit recognition result by inputting the first prompt word into the first language model. In some embodiments, the task description also indicates an inference process for outputting the explicit recognition result, and the examples further include examples of inference processes for the explicit recognition result examples.

[0063] Figure 4 illustrates a schematic diagram of an example 400 for generating prompt words that cause a language model to output explicit recognition results, according to some embodiments of the present disclosure. As shown in Figure 4, the prompt word generation module 402 may obtain page information 404 of the runtime page, a task description 406 instructing the language model to perform a task, and an example 408 including an input example 410 and an output example 412. Then, the prompt word generation module 402 may generate prompt words 418 for the language model (e.g., the first language model 312 in Figure 3) based on the page information 404, the task description 406, and the example 408.

[0064] In some embodiments, the task description 406 may include a role description that informs the language model to identify whether explicit data acquisition semantics exist in the given text, informing the user that target data has been acquired and serves a specific purpose, and to output whether the given text is runtime data acquisition explicit based on the identification result. In some embodiments, the task description 406 may also include a skill description that specifies that when the given text includes explicit data acquisition semantics, the output should be "yes," and when the given text does not include explicit data acquisition semantics, the output should be "no."

[0065] In some embodiments, task description 406 may also require the language model to output a detailed inference process. By forcing the language model to output a detailed inference process, the inference process can be made more logical, thereby improving the accuracy of the output explicit recognition results. Furthermore, requiring the language model to output a detailed inference process facilitates subsequent diagnosis and verification of the output explicit recognition results. An example of task description 406 is shown below:

[0066] "#Role

[0067] You are a professional data acquisition explicit identification and understanding expert. You can accurately identify whether a given text contains explicit data acquisition semantics that inform the user that target data has been acquired and serves a specific purpose. Based on the identification results, you can output whether the given text is a runtime data acquisition explicit, as well as the corresponding inference process.

[0068] #Skill

[0069] #Skill 1: Identifying data to obtain explicit semantics

[0070] 1. When a text message is received, it is necessary to analyze whether the text clearly informs the user that the target data has been acquired and is used for a specific purpose.

[0071] 2. If the text meets the explicit definition of data acquisition, output "Yes" and explain the reasoning process in detail.

[0072] 3. If the text does not conform to the explicit definition of data acquisition, output "No" and explain the reasoning process in detail.

[0073] #Skill 2: Interpreting explicit user meaning

[0074] 1. When a user has questions about the judgment result of a certain text, it is necessary to explain why the text conforms to or does not conform to the explicit definition of data acquisition.

[0075] 2. Provide a detailed reasoning process, explaining which parts of the text support your judgment.

[0076] As shown in Figure 4, Example 408 includes Input Example 410 and Output Example 412, where Output Example 412 may include Explicit Recognition Result Example 414. In the embodiment where the task description requires the language model to output an inference process, Output Example 412 may also include Inference Process Example 416. Example 408 helps the language model understand the task to be performed more accurately. Furthermore, Example 408 can further restrict the output format of the language model to facilitate automated processing in subsequent processes. An example of Example 408 is shown below:

[0077] #Example

[0078] Input: We obtain data A to provide you with service B. This data A will only be used for service B.

[0079] Output: Yes. Inference process: The semantics of this text are to inform the user that the mobile application will obtain data A in order to realize the function of service B. This semantics is consistent with the explicit definition of data acquisition. Furthermore, "data A is only used for service B" also informs the user that the behavior of obtaining the target data is only used to provide the user with a specific service. Therefore, the result is "Yes".

[0080] Input: This application requests data A. Do you agree or refuse?

[0081] Output: Yes. Inference process: The semantics of this text tell the user that the mobile application will obtain data A, so the result is "Yes".

[0082] Input: Data A, Data B, What is your Data C?

[0083] Output: No. Inference process: The semantics of this text do not inform the user of any data acquisition or processing actions taken by the mobile application, nor does it include any related information; therefore, it is not an explicit indication of data acquisition.

[0084] By generating the prompt word 418 based on task description 406 and example 408, the accuracy of the explicit recognition results generated by the language model for page information 404 can be improved. Furthermore, example 408 can restrict the output format of the language model to facilitate automated processing in subsequent steps. In addition, by requiring the language model to output the inference process in prompt word 418, the accuracy of the explicit recognition results can be further improved, and this also helps in the subsequent diagnosis and verification of the explicit recognition results.

[0085] In some embodiments, when generating a second prompt word for a second language model (e.g., second prompt word 316 in FIG. 3), the computing device may acquire a second task description, which indicates the determination of explicit data acquisition components included in the provided page information. The computing device may also acquire a second example, which includes an example of page information and an example of explicit data acquisition components. The computing device may then generate the second prompt word based on the page information, the second task description, and the second example. The computing device can generate the explicit data acquisition components included in the page information by inputting the second prompt word into the second language model. The computing device may then determine whether the target data meets predetermined requirements based on the explicit data acquisition components.

[0086] Figure 5 illustrates a schematic diagram of an example 500 for generating prompt words that cause a language model to output explicit components, according to some embodiments of the present disclosure. As shown in Figure 5, the prompt word generation module 502 may obtain page information 504 of the runtime page, a task description 506 instructing the language model to perform a task, and an example 508 including an input example 510 and an output example 512, wherein the output example 512 includes an explicit component example 514. The prompt word generation module 502 may then generate a prompt word 516 for a language model (e.g., the second language model 318 in Figure 3) based on the page information 504, the task description 506, and the example 508.

[0087] In some embodiments, task description 506 may include a background description, which may include definitions of explicit components. The background description helps the language model understand the meaning of explicit components more accurately, thereby improving the accuracy of the output explicit components. In some embodiments, task description 506 may include a role description, which informs the language model that it can identify explicit components present in given data and extract them, outputting them in dictionary format. In some embodiments, task description 506 may also include a skill description, which informs the language model that it possesses the skills to extract explicit components and process incomplete information. An example of task description 506 is shown below:

[0088] "#background

[0089] Data acquisition disclosure refers to the notification messages in mobile applications that inform users of the data processing actions the application will take on target data, including acquisition, sharing, and storage. The content of data acquisition disclosure includes the following seven aspects: the identity of the data controller initiating the target data processing action (IC); the user's rights in this regard (UR); the target data type being processed (TD); the purpose of the processing (PP); the legal basis for processing the target data (LB); the storage period of the processed target data (SP); and who the recipients of the processed target data are (ER).

[0090] #Role

[0091] You are a professional data acquisition explicit analysis expert who can accurately understand which explicit components exist in a given data acquisition explicit and extract them as output in dictionary format.

[0092] #Skill 1: Extracting Explicit Ingredients

[0093] 1. When providing a piece of data to obtain explicit text, it is necessary to analyze it and extract the explicit components.

[0094] 2. Output the extracted results in dictionary format, with the key-value pair settings as follows:

[0095] 'IC': The identity of the data controller who initiates the target data processing action;

[0096] 'TD': The type of target data to be processed;

[0097] 'PP': Purpose of processing;

[0098] 'UR': The rights that users enjoy regarding this;

[0099] 'LB': The legal basis for processing target data;

[0100] 'SP': Storage period for the target data being processed;

[0101] 'ER': What are the recipients of the target data being processed?

[0102] #Skill 2: Processing Incomplete Information

[0103] 1. If some explicit components are missing in the text, return null.

[0104] As shown in Figure 5, Example 508 includes Input Example 510 and Output Example 512, where Output Example 512 includes Explicit Component Example 514, the format of which corresponds to the format requirements for the explicit components of the output in Task Description 506. Below is an example of Example 508:

[0105] Input: Application A retrieves data B in order to provide you with services C.

[0106] Output: {'IC': 'Application A', 'TD': 'Data B', 'PP': 'So that we can provide you with services C', 'UR': None, 'LB': None, 'SP': None, 'ER': None}

[0107] Input: Allow application A to access data B? This will be used to improve the experience of service C. You can revoke this permission at any time. Do not allow access to the home inbox.

[0108] Output: {'IC': 'Application A', 'TD': 'Data B', 'PP': 'Used to improve the experience of service C', 'UR': 'You can revoke this permission at any time', 'LB': None, 'SP': None, 'ER': None}

[0109] Input: We need to access your camera to scan the QR code.

[0110] Output: {'IC': 'We', 'TD': 'Your Camera', 'PP': 'Scan QR Code', 'UR': None, 'LB': None, 'SP': None, 'ER': None}

[0111] By generating the cue word 516 based on the task description 506 and example 508, the accuracy of the explicit components generated by the language model can be improved. Furthermore, example 508 can restrict the output format of the language model to facilitate automated processing in subsequent steps.

[0112] In some embodiments, during the training phase of a first language model (e.g., first language model 312 in FIG. 3), a computing device may acquire multiple page information from multiple applications and multiple tags for the multiple page information, each of the multiple tags indicating whether the corresponding page information in the multiple page information includes explicit data acquisition. The computing device then generates a language model by training a first pre-trained language model based on the multiple page information and the multiple tags.

[0113] Figure 6 illustrates a schematic diagram of an example 600 for training a language model to output explicit data acquisition results according to some embodiments of the present disclosure. As shown in Figure 6, in example 600, the language model 610 can be a pre-trained language model such as GPT or BERT, which possesses powerful language understanding and generation capabilities through large-scale pre-training. However, pre-trained language models have a broad range of knowledge and skills, but lack domain-specific expertise, accuracy, and task adaptability. Therefore, in example 600, the language model 610 can be fine-tuned based on a training dataset 608 to enable the language model 610 to learn the expertise regarding explicit data acquisition in the training dataset 608. After training, the language model 610 can accurately identify whether the page information contains explicit data acquisition and output explicit data acquisition results. The trained language model 610 can be, for example, the first language model 312 in Figure 3.

[0114] In Example 600, the computing device can acquire multiple page information 604-1, 604-2, ..., and 604-N (collectively referred to as page information 604) from multiple runtime pages of multiple applications 602-1, 602-2, ..., and 602-N (collectively referred to as application 602). Then, the computing device can acquire multiple truth-value explicit identification results 606-1, 606-2, ..., and 606-N (collectively referred to as truth-value explicit identification results 606) for the multiple page information 604. For example, the truth-value explicit identification result 606 can be generated manually based on the page information 604. Thus, the page information 604 and the truth-value explicit identification result 606 can constitute a training dataset 608.

[0115] In Example 600, language model 610 can generate a predicted explicit recognition result based on page information 604. For example, language model 610 can generate a predicted explicit recognition result 612 based on page information 604-N. The computing device can then compare the predicted explicit recognition result 612 with the ground truth explicit recognition result 606-N of page information 604-N to generate a loss 614. Loss 614 can indicate the difference between the predicted explicit result 612 and the ground truth explicit result 606-N. The computing device can then use loss 614 to fine-tune language model 610, thereby generating a trained language model 610.

[0116] By performing supervised fine-tuning on a pre-trained language model, its existing semantic understanding capabilities can be leveraged. Furthermore, the language model can learn expertise regarding explicit data retrieval from the training dataset 608, thereby improving the accuracy of explicit data recognition.

[0117] In some embodiments, during the training of the second language model, the computing device may acquire multiple claim documents displayed in multiple applications to describe the target data. The computing device may also extract multiple data-derived explicit components from the multiple claim documents. The computing device can then generate the second language model by training a second pre-trained language model based on the multiple data-derived explicit components.

[0118] Figure 7 illustrates a schematic diagram of an example 700 for training a language model for outputting explicit components according to some embodiments of the present disclosure. In the runtime pages of an application, the explicit content of data acquisition is often scattered; however, the semantics of the explicit components are often fixed and consistent with the semantics of the corresponding content in the declaration documents describing the target data (e.g., privacy policy documents). For example, explicit components such as the target data type and acquisition purpose displayed on the runtime pages also have corresponding expressions in the declaration documents, and they are semantically consistent. Therefore, in embodiments of the present disclosure, the pre-trained language model can be fine-tuned by transferring knowledge from the declaration documents in a few-shot learning manner, by drawing on the explicit component expectations from a large number of easily accessible declaration documents.

[0119] In Example 700, language model 710 is a pre-trained language model. Language model 710 can be language model 610 in Figure 6 or another language model. The computing device can acquire multiple declaration documents 704-1, 704-2, ..., and 704-N (collectively referred to as declaration documents 704) for multiple applications 702-1, 702-2, ..., and 702-N (collectively referred to as application 702). In addition, the computing device can acquire multiple pre-identified explicit components 706-1, 706-2, ..., and 706-N (collectively referred to as explicit components 706) from the multiple declaration documents 704. Thus, the declaration documents 704 and the pre-identified explicit components 706 can constitute a training dataset 708.

[0120] In Example 700, the computing device can fine-tune the language model 710 based on the training dataset 708, so that the language model 710 learns the knowledge about the explicit component 706 in the declaration document 704. In this way, the problem of lack of explicit component corpus caused by the difficulty in obtaining explicit data at runtime can be solved.

[0121] Figure 8 shows a block diagram of an apparatus 800 for detecting explicit data acquisition in an application according to some embodiments of the present disclosure. As shown in Figure 8, the apparatus 800 includes an acquisition behavior determination module 802 configured to determine that the application acquires target data. The apparatus 800 also includes a page information acquisition module 804 configured to acquire page information on a runtime page in the application. The apparatus 800 also includes an explicit identification module 806 configured to generate an explicit identification result based on the page information using a language model, the explicit identification result indicating whether the page information includes explicit data acquisition. Furthermore, the apparatus 800 includes an explicit detection module 808 configured to determine whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit identification result.

[0122] In some embodiments, the behavior determination module 802 includes: a system interface acquisition module configured to acquire system interfaces called by an application; and a system interface analysis module configured to determine that the application acquires target data by determining that the system interface is used to acquire target data; or a network data acquisition module configured to acquire network data sent by the application; and a network data analysis module configured to determine that the application acquires target data based on determining that the network data includes target data.

[0123] In some embodiments, the page information acquisition module 804 includes: a single sign-on module configured to log in to the application by logging in to a third-party account and using the single sign-on method; and a page information acquisition submodule configured to acquire page information on the runtime page of the application while logged in.

[0124] In some embodiments, the explicit identification module 806 includes: a first task description acquisition module configured to acquire a first task description, the first task description indicating whether there is explicit data acquisition information in the provided page information that informs the user that target data has been acquired and that the target data is used for a specific purpose; a first example acquisition module configured to acquire a first example, the example including a page information example and an explicit identification result example; a first prompt word generation module configured to generate a first prompt word based on the page information, the first task description, and the first example; and a first prompt word usage module configured to generate an explicit identification result by inputting the first prompt word into a language model.

[0125] In some embodiments, the task description further indicates an inference process for outputting explicit identification results, and the examples also include examples of inference processes for examples of explicit identification results.

[0126] In some embodiments, the explicit detection module 808 includes: an association determination module configured to determine that the page information is associated with target data in response to an explicit identification result indicating that the page information includes explicit data acquisition; a first time determination module configured to determine the acquisition time of the target data; a second time determination module configured to determine the display time of the runtime page; and a time comparison module configured to determine that the acquisition of the target data does not meet predetermined requirements in response to the acquisition time of the target data being earlier than the display time of the runtime page.

[0127] In some embodiments, the association determination module includes: a first data category determination module configured to determine a first data category of the target data; a second data category determination module configured to determine a second data category associated with the page information; a semantic relevance determination module configured to determine the semantic relevance between the first data category and the second data category, wherein the semantic relevance indicates whether the first data category and the second data category are semantically similar or whether the first data category and the second data category are semantically hierarchical; and a semantic relevance usage module configured to determine the association between the target data and the page information based on the semantic relevance.

[0128] In some embodiments, where the language model is a first language model, the apparatus 800 further includes a second language model using module, configured to, in response to determining that page information is associated with target data, use the second language model based on the page information to determine whether the acquisition of the target data meets predetermined requirements.

[0129] In some embodiments, the second language model using module, based on page information, includes: a second task description acquisition module configured to acquire a second task description, the second task description indicating the explicit data acquisition components included in the provided page information; a second example acquisition module configured to acquire a second example, the second example including a page information example and an example of explicit data acquisition components; a second prompt word generation module configured to generate a second prompt word based on the page information, the second task description, and the second example; a second prompt word using module configured to generate the explicit data acquisition components included in the page information by inputting the second prompt word into the second language model; and an explicit component using module configured to determine whether the target data meets predetermined requirements based on the explicit data acquisition components.

[0130] In some embodiments, the explicit component usage module includes an explicit component analysis module configured to determine that the acquisition of target data does not meet the predetermined requirements in response to the absence of an explicit component specified in the predetermined requirements for data acquisition.

[0131] In some embodiments, the apparatus 800 further includes: a declaration document acquisition module configured to acquire multiple declaration documents displayed in multiple applications for describing target data; a declaration document usage module configured to extract multiple data acquisition explicit components from the multiple declaration documents; and a second language model training module configured to generate a second language model by training a second pre-trained language model based on the multiple data acquisition explicit components.

[0132] In some embodiments, the apparatus 800 further includes: a tag acquisition module configured to acquire multiple page information from multiple applications and multiple tags for the multiple page information, each of the multiple tags indicating whether the corresponding page information in the multiple page information includes explicit data acquisition; and a first language model training module configured to generate a language model by training a first pre-trained language model based on the multiple page information and the multiple tags.

[0133] It is understood that by utilizing the apparatus 800 of this disclosure, at least one of the many advantages achievable by the methods or processes described above can be realized. For example, the language model can identify whether the page information displayed on the runtime page is explicit data acquisition, reducing false and false identifications, thereby improving the accuracy of identifying explicit data acquisition. Furthermore, by detecting data acquisition behavior in the application, it is possible to determine whether the data acquisition behavior meets predetermined requirements based on the explicit identification results of the runtime page, thereby improving the comprehensiveness of explicit data acquisition detection.

[0134] Figure 9 shows a block diagram of a device 900 capable of implementing various embodiments of the present disclosure. Device 900 may, for example, be a computing device 102 as shown in Figure 1. As shown in Figure 9, device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. Various programs and data required for the operation of device 900 may also be stored in RAM 903. The CPU / GPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904. Although not shown in Figure 9, device 900 may also include a coprocessor.

[0135] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0136] The various methods or processes described above can be executed by CPU / GPU 901. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU / GPU 901, one or more steps or actions in the methods or processes described above can be performed.

[0137] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0138] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0139] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0140] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0141] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0142] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0144] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting explicit data acquisition in an application, comprising: The application is determined to acquire the target data; Obtain page information from the runtime page in the application; Based on the page information, a language model is used to generate an explicit recognition result, which indicates whether the page information includes explicit data acquisition information. as well as Based on the target data and the explicit identification result, it is determined whether the acquisition of the target data meets the predetermined requirements.

2. The method according to claim 1, wherein determining that the application acquires the target data includes: Obtain the system interface called by the application; as well as The system interface is used to obtain the target data to determine whether the application obtains the target data. Obtain network data sent by the application; as well as The application determines to acquire the target data based on the determination that the network data includes the target data.

3. The method according to claim 1, wherein obtaining the page information on the runtime page in the application includes: Log in to the application by logging into a third-party account and using single sign-on; as well as Obtain the page information on the runtime page of the application while logged in.

4. The method according to claim 1, wherein generating the explicit recognition result using the language model based on the page information includes: Obtain a first task description, which indicates whether there is an explicit data acquisition statement in the provided page information that informs the user that target data has been acquired and that the target data is used for a specific purpose. Obtain a first example, which includes a page information example and an explicit recognition result example; The first prompt word is generated based on the page information, the first task description, and the first example; as well as The explicit recognition result is generated by inputting the first prompt word into the language model.

5. The method of claim 4, wherein the task description further indicates an inference process for outputting explicit identification results, and the example further includes an example of an inference process for the example of the explicit identification results.

6. The method according to claim 1, wherein determining whether the acquisition of the target data meets the predetermined requirements based on the target data and the explicit identification result includes: In response to the explicit identification result indicating that the page information includes the explicit data acquisition, it is determined that the page information is associated with the target data; Determine the acquisition time of the target data; Determine the display time of the runtime page; as well as In response to the fact that the acquisition time of the target data is earlier than the display time of the runtime page, it is determined that the acquisition of the target data does not meet the predetermined requirements.

7. The method of claim 6, wherein determining that the target data is associated with the page information comprises: Determine the first data category of the target data; Determine a second data category associated with the page information; Determine the semantic relevance between the first data category and the second data category. The semantic relevance indicates whether the first data category and the second data category are semantically similar or whether the first data category and the second data category are semantically hierarchical. as well as The semantic relevance is used to determine the association between the target data and the page information.

8. The method according to claim 6, wherein the language model is a first language model, and the method further comprises: In response to determining that the page information is associated with the target data, a second language model is used based on the page information to determine whether the acquisition of the target data meets the predetermined requirements.

9. The method according to claim 8, wherein determining whether the acquisition of the target data meets the predetermined requirements based on the page information using the second language model includes: Obtain a second task description, which indicates the determination of explicit data acquisition components included in the provided page information; Obtain a second example, which includes a page information example and a data acquisition explicit component example; The second prompt word is generated based on the page information, the second task description, and the second example; The explicit components of the page information are obtained by inputting the second prompt word into the second language model. as well as Based on the data, the explicit components are obtained to determine whether the target data meets the predetermined requirements.

10. The method of claim 9, wherein determining whether the acquisition of the target data meets the predetermined requirements based on the data acquisition explicit components comprises: In response to the absence of the explicit component specified in the predetermined requirement in the data acquisition explicit component, it is determined that the acquisition of the target data does not meet the predetermined requirement.

11. The method of claim 9, further comprising: Retrieve multiple claim documents displayed in multiple applications to describe the target data; Extract multiple data points from the multiple declaration documents to obtain explicit components; as well as The second language model is generated by training a second pre-trained language model based on explicit components obtained from the multiple data sets.

12. The method according to claim 1, further comprising: Obtain multiple page information from multiple applications and multiple tags for the multiple page information, each of the multiple tags indicating whether the corresponding page information in the multiple page information includes explicit data acquisition; as well as The language model is generated by training a first pre-trained language model based on the multiple page information and the multiple tags.

13. An apparatus for detecting explicit data acquisition in an application, comprising: The behavior determination module is configured to determine the target data that the application acquires; The page information acquisition module is configured to acquire page information on the runtime pages in the application; The explicit identification module is configured to generate an explicit identification result based on the page information using a language model, wherein the explicit identification result indicates whether the page information includes explicit data acquisition information; as well as The explicit detection module is configured to determine whether the acquisition of the target data meets predetermined requirements based on the target data and the explicit identification result.

14. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 12.

15. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to any one of claims 1 to 12.