Privacy policy legality automatic detection method and system based on deep learning

By employing a deep learning-based approach, combined with multi-dimensional corpus screening and a dual-task architecture, the accuracy and efficiency issues of legality detection for Chinese App policy texts have been addressed. This approach enables efficient and interpretable compliance detection, applicable to various industry scenarios.

CN120929919APending Publication Date: 2025-11-11CENT SOUTH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511052946.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies are insufficient to efficiently and accurately determine the legality of privacy policies in Chinese App policy texts, and lack interpretability and stability in terms of legal applicability, thus failing to meet the requirements for accuracy and consistency in compliance testing.

Method used

We employ a deep learning-based approach, using a multi-dimensional corpus screening mechanism driven by industry risk weights to construct fine-grained semantic judgment rules. We utilize a large language model to generate legality labeling results and combine BERT, CNN, and BiLSTM models for legality classification and violation type identification, forming a dual-task architecture that supports multi-dimensional compliance detection.

Benefits of technology

It enables structured automatic parsing and intelligent legality recognition of privacy policy texts, improving the accuracy and efficiency of detection, possessing cross-industry adaptability, supporting legal risk governance and compliance management, and providing interpretable compliance judgment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929919A_ABST
    Figure CN120929919A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy policy legality automatic detection method and system based on deep learning, which is applied to the technical field of legal artificial intelligence, and comprises the following steps: obtaining a privacy policy text based on a multi-dimensional corpus screening mechanism driven by an industry risk weight; on the basis of the determined compliance dimension and the legality state label, constructing a fine-grained semantic judgment rule, and generating a legality labeling result by using a large language model; constructing a dual-task architecture of legality classification and violation type identification based on a deep learning language model, and performing model training based on the obtained legality annotation data set; according to the dual-task architecture, BERT is used as a basic coding layer, and a CNN and a BiLSTM are respectively adopted to carry out legality classification and violation type identification; and inputting a to-be-tested privacy policy text to the trained model to obtain an identification result. According to the method, the limitation of traditional rule matching is broken through, and automatic compliance analysis, multi-dimensional risk positioning and interpretability judgment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of legal artificial intelligence technology, and more specifically to a method and system for automatically detecting the legality of privacy policies based on deep learning. Background Technology

[0002] With the rapid development of mobile internet and artificial intelligence technologies, apps are widely used in various scenarios such as social networking, travel, shopping, and healthcare. The collection and use of users' personal information has become a typical, high-frequency processing activity. To strengthen the regulation of data processing activities, more explicit and mandatory compliance requirements have been imposed on app operators. However, privacy policies, as the primary means for app operators to disclose data processing activities, serve a dual function as both "basis for user authorization" and "certificate of regulatory compliance." However, actual privacy policies often suffer from problems such as excessive length, disorganized structure, and abstract expression, making them difficult for ordinary users to understand, and posing semantic challenges for administrative and judicial review.

[0003] Currently, compliance reviews of privacy policies mainly rely on manual verification or rule matching, which cannot accurately map to the stipulated principles such as clear processing purpose, minimum necessity, and individual consent. Furthermore, the lack of unified review standards makes them prone to subjective bias, and compliance judgments are difficult to quantify and accept. Therefore, there is an urgent need for more systematic automated review methods as technical support.

[0004] Currently, three main approaches to automated compliance detection have emerged, but all have certain limitations: First, NLP-based semantic analysis algorithms. Tools like Polisis, based on semantic segmentation and rule templates, attempt to analyze privacy policies for compliance. While they can achieve paragraph-level legality classification and data collection purpose identification, they often remain at the surface-level syntactic matching level, failing to accurately determine whether legal provisions are substantively satisfied, and exhibiting a disconnect from the legal logic behind the privacy policy. Second, deep learning model-based detection. Some studies construct privacy policy data as a training set for classification tasks, employing pre-trained language models (such as BERT) + CNN / BiLSTM architectures to achieve "legal / illegal" or "multi-label violation type identification." However, since the label system is mostly empirically set and lacks a clear mapping to the legal structure, while the model can classify text, it cannot explain the basis for judgment, lacking interpretability and stability in terms of legal application. Third, detection methods combining static and dynamic analysis. In some practices, user behavior data and data call logs have been incorporated to form a multimodal compliance analysis system combining text and behavior, enabling more accurate data risk prediction. However, such solutions rely on backend permissions, have high deployment costs, and cannot address the issue of unclear expression of the legality of policy texts themselves.

[0005] It is evident that current technology lacks a legality detection method that is centered on legal provisions, possesses structural mapping capabilities, and can output the reasons for compliance judgments. This is especially true in Chinese app policy texts, where the expression of compliance risks is highly ambiguous and flexible.

[0006] Therefore, how to provide a deep learning-based automatic detection method and system for privacy policy legality that integrates legal rule semantics, intelligent annotation system and highly interpretable model structure is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a method and system for automatically detecting the legality of privacy policies based on deep learning.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A deep learning-based method for automatically detecting the legality of privacy policies includes:

[0010] Step 1: Based on the industry risk weight-driven multi-dimensional corpus screening mechanism, obtain the privacy policy text, and construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and use a large language model to generate legality labeling results;

[0011] Step 2: Construct a dual-task architecture for legality classification and violation type identification based on a deep learning language model, and train the model based on the legality annotation dataset obtained in Step 1; wherein, the dual-task architecture uses the pre-trained language model BERT as the basic encoding layer, and uses a convolutional neural network CNN and a bidirectional long short-term memory network BiLSTM for legality classification and violation type identification respectively.

[0012] Step 3: Input the privacy policy text to be tested into the model trained in Step 2 to obtain the results of legality classification and violation type identification.

[0013] Optionally, in step 1, the privacy policy text is obtained based on a multi-dimensional corpus screening mechanism driven by industry risk weights, specifically as follows:

[0014] We set sampling weights based on multiple industry risk indicators, including the App's industry category, monthly active users, app download popularity, and the sensitivity level of the data involved in the service, and obtained the privacy policy text.

[0015] Optionally, in step 1, after obtaining the privacy policy text, the following steps are also included: using a composite cleaning strategy that combines regular expressions and syntax boundary recognition to perform multiple preprocessing tasks on the privacy policy text, including webpage tag stripping, footnote redundancy removal, invalid link filtering, empty segment identification and removal, and duplicate sentence compression.

[0016] Optionally, in step 1, after performing multiple preprocessing tasks on the privacy policy text, the following steps are also included: using a paragraph-statement two-level indexing system to perform structured encoding on the preprocessed privacy policy text, retaining the original title number, clause sequence and contextual structure information in each statement.

[0017] Optionally, in step 1, fine-grained semantic judgment rules are constructed based on the determined compliance dimensions and the legality status labels under the compliance dimensions, and legality annotation results are generated using a large language model, specifically as follows:

[0018] Based on the six compliance dimensions and four levels of legal status labels, fine-grained semantic judgment rules are constructed for each compliance dimension. The compliance dimensions include: informed consent, openness and transparency, individual consent, necessity of processing, clear purpose, and correction and deletion. The four levels of legal status labels include: not involved, legal, potentially illegal, and obviously illegal.

[0019] Based on fine-grained semantic judgment rules, corresponding legal instructions and contextual examples are constructed for each compliance dimension to form a structured prompt template, and legality labeling results are generated through the large language model ChatLaw.

[0020] Optionally, in step 1, after generating the legality labeling results using the large language model, the following steps are also included: introducing a manual review mechanism and a label consistency logic check.

[0021] Optionally, step 2 also includes: introducing an attention mechanism based on the violation type identification using the Bi-directional Long Short-Term Memory (BiLSTM) network to guide the model to focus on key legal terms and behavioral expressions.

[0022] Optionally, step 2 may also include: during model training, using a cross-entropy loss function for the legality classification task and a binary cross-entropy loss function for the violation type identification task.

[0023] Optionally, step 2 also includes: during model training, selecting the AdamW optimizer and combining it with a warmup mechanism, while introducing a Dropout regularization strategy in both the Convolutional Neural Network (CNN) and the Bidirectional Long Short-Term Memory (BiLSTM) network.

[0024] This invention also provides a deep learning-based automatic privacy policy legality detection system that utilizes a deep learning-based automatic privacy policy legality detection method, comprising:

[0025] Text Acquisition and Annotation Module: Used to acquire privacy policy texts based on a multi-dimensional corpus screening mechanism driven by industry risk weights, and to construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and to generate legality annotation results using a large language model;

[0026] Model building and training module: This module is used to build a dual-task architecture for legality classification and violation type recognition based on a deep learning language model, and to train the model based on the legality annotation dataset obtained in the text acquisition and annotation module. The dual-task architecture uses the pre-trained language model BERT as the base encoding layer, and employs a convolutional neural network (CNN) and a bidirectional long short-term memory network (BiLSTM) for legality classification and violation type recognition, respectively.

[0027] Legality identification module: This module is used to input the privacy policy text to be tested into the model trained in the model building and training module, and obtain the legality classification and violation type identification results.

[0028] As can be seen from the above technical solution, compared with the prior art, this invention discloses a method and system for automatic detection of privacy policy legality based on deep learning. In the task of privacy policy legality detection, this invention breaks through the performance bottlenecks of traditional rule matching and shallow models. Around core aspects such as data construction, annotation system, model architecture, and system deployment, it forms a technically complete, structurally clear, and highly practical solution. Its specific advantages are reflected in the following four aspects:

[0029] 1. Significantly improved review efficiency:

[0030] This invention achieves structured automatic parsing and intelligent legality recognition of privacy policy texts, enabling efficient processing of large-scale corpora, significantly reducing manual intervention, and improving review efficiency. The annotation stage incorporates a large language model to generate preliminary judgments, combined with manual review and consistency checks, reducing the cost of acquiring high-quality supervisory data.

[0031] 2. Enhanced detection accuracy and judgment capability:

[0032] By combining the contextual understanding capabilities of the BERT model with the multi-layer feature extraction mechanism of the CNN / BiLSTM structure, this invention supports the accurate identification of complex sentence structures, ambiguous expressions, and multi-label risks. Based on the core provisions of protection laws and the regulatory-oriented labeling system built upon relevant practical standards, the model exhibits higher robustness and interpretability in legality determination and dimensional classification.

[0033] 3. The system has strong adaptability and scalability:

[0034] This invention employs a modular system structure and a flexible tagging mechanism, supporting rapid expansion of compliance dimensions and model structure, and adapting to different national legal systems and industry standards. The system already possesses practical implementation capabilities for high-risk scenarios such as finance, e-commerce, and social media, and has the foundation for cross-industry promotion.

[0035] 4. Supporting institutionalized compliance management and legal risk governance:

[0036] Through the mapping of regulatory provisions and the structured output mechanism, this invention can not only serve as an auxiliary tool for enterprise compliance review but also provide quantitative assessment support for regulatory platforms. The model's judgment results have clear boundaries and strong label interpretability, which can assist legal professionals in classifying violations and analyzing the basis for accountability, thus promoting the institutionalization of personal information protection governance. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0038] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0039] Figure 2 This is a schematic diagram of the dual-task architecture of legality classification and violation type identification based on a deep learning language model according to the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1:

[0042] Embodiment 1 of this invention discloses an automatic detection method for the legality of privacy policies based on deep learning, such as... Figure 1 As shown, it includes:

[0043] Step 1: Based on the industry risk weight-driven multi-dimensional corpus screening mechanism, obtain the privacy policy text, and construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and use a large language model to generate legality labeling results.

[0044] To address the challenges of insufficient industry representativeness, inconsistent data quality, and sparse compliance features in current Chinese privacy policy text data, which make it difficult to support high-precision model training and multi-dimensional label representation, this invention designs a corpus construction method for privacy policy legality identification tasks. The aim is to build a data foundation with high representativeness, high quality, and high compliance relevance by introducing industry sensitivity indicators and regulatory constraints.

[0045] In the corpus collection stage, this invention proposes a multi-dimensional corpus screening mechanism driven by industry risk weights. This mechanism uses mainstream mobile application stores as data crawling sources. Through this mechanism, a Chinese privacy policy dataset with diverse semantic content, balanced industry distribution, and dense compliance elements is finally constructed.

[0046] Based on a multi-dimensional corpus screening mechanism driven by industry risk weights, the privacy policy text was obtained, specifically:

[0047] We set sampling weights based on multiple industry risk indicators, including the App's industry category, monthly active users, app download popularity, and the sensitivity level of the data involved in the service, and obtained the privacy policy text.

[0048] Furthermore, priority is given to high-risk industries that rely heavily on personal information, such as finance, e-commerce, social media, and education.

[0049] In the data cleaning stage, to address issues such as nested HTML tags, disordered paragraph structure, and inconsistent format encoding in the original policy documents, this invention introduces a composite cleaning strategy that combines regular expressions with grammatical boundary recognition. This ensures that the corpus retains complete legal semantics while improving its annotability and subsequent modeling quality.

[0050] Therefore, after obtaining the privacy policy text, the process also includes: using a composite cleaning strategy that combines regular expressions and syntax boundary recognition to perform multiple preprocessing tasks on the privacy policy text, including stripping web page tags, removing redundant footnotes, filtering invalid links, identifying and removing empty segments, and compressing duplicate sentences.

[0051] Furthermore, to better support subsequent tasks such as label system construction and compliance determination, this invention also designs a legally friendly structured coding method. By combining paragraph logic and syntactic boundaries, this method constructs a two-layer "paragraph-statement" indexing system to enhance the model's ability to identify hierarchical logic and semantic dependencies of regulations, ultimately providing structural support and a legal context mapping foundation for legality dimension identification.

[0052] Therefore, after performing multiple preprocessing tasks on the privacy policy text, the process also includes: using a paragraph-statement two-level indexing system to structurally encode the preprocessed privacy policy text, retaining the original title number, clause sequence, and contextual structure information in each statement.

[0053] To achieve structured modeling and compliance judgment of legality elements in privacy policy texts, this invention designs a multi-stage annotation system that integrates legal normative knowledge and the prompting capabilities of a large language model. This system not only performs dimensional analysis and hierarchical classification of privacy policy texts, but also constructs a four-stage closed-loop annotation mechanism of "prompt-response-review-verification" to ensure that semantic tags are both legally grounded and algorithmically compatible.

[0054] Based on the core provisions of the protection law and relevant practical standards, the legality requirements for personal information processing in privacy policies were abstracted and structurally transformed. Six compliance dimensions were ultimately identified: informed consent, openness and transparency, individual consent, necessity of processing, clear purpose, and correction and deletion. Furthermore, a four-level legality status label was established: 0 - not involved, 1 - legal, 2 - potentially illegal, and 3 - obviously illegal.

[0055] Based on the determined compliance dimensions and the legality status labels under those dimensions, fine-grained semantic judgment rules are constructed, and legality annotation results are generated using a large language model, specifically:

[0056] Based on the six compliance dimensions and four levels of legal status labels, fine-grained semantic judgment rules are constructed for each compliance dimension to clarify the boundaries between legality and illegality. The compliance dimensions include: informed consent, openness and transparency, individual consent, necessity of processing, clear purpose, and correction and deletion. The four levels of legal status labels include: not involved, legal, potentially illegal, and obviously illegal.

[0057] To improve annotation efficiency and reduce manual workload, this invention introduces the large language model ChatLaw and designs a prompt-based label generation process. Based on fine-grained semantic judgment rules, corresponding legal instructions and contextual examples are constructed for each compliance dimension to form a structured prompt template. The ChatLaw large language model generates legality annotation results, which have advantages such as strong context adaptability and strong generalization ability of compliance language.

[0058] After generating the legality labeling results using a large language model, the process also includes: introducing a manual review mechanism and a label consistency logic check.

[0059] For the manual review mechanism, all labeling results must be independently judged by two reviewers with legal backgrounds. Samples with conflicting or inconsistent results are submitted to legal experts for final determination.

[0060] For label consistency logic verification, the system's built-in module can automatically identify logical conflict situations such as "a statement lacking a purpose description is labeled as legal," thereby correcting label errors.

[0061] To complement the annotation process, this invention also developed a dedicated semantic annotation tool that supports functions such as text structure segmentation, label annotation, multi-round review, consistency feedback, and annotation record tracking. It can export annotation results in a standardized JSON format, supporting direct access to model training and compliance review platforms. This system ensures alignment between compliance labels and legal regulations while achieving efficient integration of intelligent annotation of large models and manual review, demonstrating good scalability and practical application value.

[0062] Step 2: Construct a dual-task architecture for legality classification and violation type identification based on a deep learning language model, such as... Figure 2 As shown, the model is trained based on the legality annotation dataset obtained in step 1; the dual-task architecture uses the pre-trained language model BERT as the basic encoding layer, and uses convolutional neural network CNN and bidirectional long short-term memory network BiLSTM for legality classification and violation type identification, respectively.

[0063] This invention focuses on the core objective of judging the compliance of privacy policy texts. It constructs a dual-task architecture based on deep language models for legality classification and violation type identification, aiming to solve three key problems simultaneously: first, the distinction between compliant clauses and potentially non-compliant expressions; second, the multi-label identification problem where a single sentence may involve multiple non-compliant dimensions; and third, how to improve the model's ability to identify key legal elements and semantic interpretation capabilities in long and complex sentences.

[0064] To address this, this invention proposes a multi-module integrated structured modeling scheme. The encoding layer uses the pre-trained language model BERT as the base network to obtain the contextual embeddings of sentences within the given context, enhancing the model's ability to understand long sentences and complex expressions in Chinese privacy policies. In the legality judgment task, a convolutional neural network (CNN) is introduced as a local feature extractor, superimposed on the BERT output layer, to extract significant phrase patterns and implicit violations, completing the legal / illegal binary classification of each sentence. To further identify the specific compliance dimensions involved in illegal sentences, this invention constructs a bidirectional long short-term memory (BiLSTM) network structure based on the shared BERT encoding to model long-distance semantic dependencies within sentences. Simultaneously, an attention mechanism is introduced to guide the model to focus on key legal terms and behavioral expressions such as "purpose," "authorization," and "subject," thereby improving the model's ability to perceive and discriminate fine-grained violation labels.

[0065] The two task structures described above run in parallel by sharing a BERT encoding layer and employ end-to-end joint training to achieve semantic transfer and task consistency optimization. This structure can simultaneously output legality judgment results and multi-dimensional violation labels, enabling the model to balance classification accuracy and compliance interpretation capabilities within a unified architecture. This makes it suitable for the dual requirements of "judgment + interpretation" in real-world privacy review scenarios.

[0066] During model training, the Cross-Entropy Loss function is used for the binary classification task of legality, while the Binary Cross-Entropy Loss function (BCELoss) is used for the multi-label recognition task of illegal labels, ensuring the accuracy and stability of the judgment results in each dimension. The AdamW optimizer is selected, with an initial learning rate of 5e-5, and a warmup mechanism is used to improve the robustness of the model in the early stages of training. The entire training process consists of 60 epochs, and Dropout regularization strategies (with a ratio of 0.1 to 0.3) are introduced into the CNN and BiLSTM modules to mitigate the risk of overfitting and improve the model's generalization ability on real data.

[0067] Meanwhile, to verify the model's recognition performance across different compliance dimensions, this invention employs multiple classification metrics such as precision, recall, and F1 score, and calculates sub-metrics for each of the six compliance label dimensions to comprehensively evaluate the model's performance balance and risk identification capabilities in multi-task scenarios.

[0068] Step 3: Input the privacy policy text to be tested into the model trained in Step 2 to obtain the results of legality classification and violation type identification.

[0069] The compliance assessment method proposed in this invention possesses excellent structural clarity and deployment adaptability, and can be widely applied to various scenarios such as government supervision, enterprise data auditing, and app store compliance review. The method consists of five stages: "text preprocessing → model judgment → result output → system integration → deployment adaptation," forming a complete end-to-end automated compliance analysis process.

[0070] In the text processing stage, the privacy policy to be detected first undergoes preprocessing operations such as syntactic cleaning, format normalization, and paragraph segmentation to ensure that the input structure meets the model's recognition requirements. Subsequently, the model performs a legality judgment on each statement, and after identifying non-compliant expressions, it further outputs the corresponding violation dimension label (e.g., "unclear purpose", "lack of separate consent", etc.), thereby achieving an interpretable and structured compliance assessment result.

[0071] The system outputs results in JSON or tabular format, facilitating integration into various front-end review platforms, enterprise review workflows, or compliance report generation systems. In regulatory scenarios, the system assists reviewers in quickly screening risky statements and marking the risk level of clauses. In enterprise applications, it can be used for real-time compliance alerts and violation identification during the drafting of new policy texts, significantly reducing manual review costs.

[0072] To meet the deployment needs of computing-constrained scenarios such as mobile devices, small application systems, or edge nodes, this invention further supports model compression and lightweight adaptation solutions. By optimizing the original model through structural pruning, knowledge distillation, and quantization compression, a lightweight compliance judgment model can be generated and deployed on local devices or embedded systems, ensuring the sustainable operation of compliance audit capabilities across multiple scenarios and terminals.

[0073] Example 2:

[0074] Embodiment 2 of this invention discloses a deep learning-based automatic privacy policy legality detection system, which utilizes a deep learning-based automatic privacy policy legality detection method, comprising:

[0075] Text Acquisition and Annotation Module: This module uses a multi-dimensional corpus screening mechanism driven by industry risk weights to acquire privacy policy texts, construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and generate legality annotation results using a large language model.

[0076] Model building and training module: This module is used to build a dual-task architecture for legality classification and violation type recognition based on a deep learning language model, and to train the model based on the legality annotation dataset obtained in the text acquisition and annotation module. The dual-task architecture uses the pre-trained language model BERT as the base encoding layer, and employs a convolutional neural network (CNN) and a bidirectional long short-term memory network (BiLSTM) for legality classification and violation type recognition, respectively.

[0077] Legality identification module: This module is used to input the privacy policy text to be tested into the model trained in the model building and training module, and obtain the legality classification and violation type identification results.

[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0079] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for automatically detecting the legality of a privacy policy based on deep learning, characterized in that, include: Step 1: Based on the industry risk weight-driven multi-dimensional corpus screening mechanism, obtain the privacy policy text, and construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and use a large language model to generate legality labeling results; Step 2: Construct a dual-task architecture for legality classification and violation type identification based on a deep learning language model, and train the model based on the legality annotation dataset obtained in Step 1; wherein, the dual-task architecture uses the pre-trained language model BERT as the basic encoding layer, and uses a convolutional neural network CNN and a bidirectional long short-term memory network BiLSTM for legality classification and violation type identification respectively. Step 3: Input the privacy policy text to be tested into the model trained in Step 2 to obtain the results of legality classification and violation type identification.

2. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, In step 1, the privacy policy text is obtained based on a multi-dimensional corpus screening mechanism driven by industry risk weights, specifically as follows: The privacy policy text was obtained by setting sampling weights based on multiple industry risk indicators, including the industry category of the app, the number of monthly active users, the app download popularity, and the sensitivity level of the data involved in the service.

3. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, Step 1, after obtaining the privacy policy text, also includes: using a composite cleaning strategy combining regular expressions and syntax boundary recognition to perform multiple preprocessing tasks on the privacy policy text, including webpage tag stripping, footnote redundancy removal, invalid link filtering, empty segment identification and removal, and duplicate sentence compression.

4. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 3, characterized in that, In step 1, after performing multiple preprocessing tasks on the privacy policy text, the following steps are also included: using a paragraph-statement two-level indexing system to perform structured encoding on the preprocessed privacy policy text, retaining the original title number, clause sequence and contextual structure information in each statement.

5. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, In step 1, fine-grained semantic judgment rules are constructed based on the determined compliance dimensions and the legality status labels under the compliance dimensions. A large language model is then used to generate legality annotation results, specifically: Based on the identified six compliance dimensions and four levels of legal status labels, fine-grained semantic judgment rules are constructed for each compliance dimension. The compliance dimensions include: informed consent, openness and transparency, individual consent, necessity of processing, clear purpose, and correction and deletion. The four levels of legal status labels include: not involved, legal, potentially illegal, and obviously illegal. Based on the fine-grained semantic judgment rules, corresponding legal instructions and contextual examples are constructed for each compliance dimension to form a structured prompt template, and legality annotation results are generated through the large language model ChatLaw.

6. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, Step 1, after generating the legality labeling results using the large language model, also includes: introducing a manual review mechanism and a label consistency logic check.

7. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, Step 2 also includes: introducing an attention mechanism based on the violation type identification performed by the Bidirectional Long Short-Term Memory Network (BiLSTM) to guide the model to focus on key legal terms and behavioral expressions.

8. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, Step 2 also includes: during model training, using the cross-entropy loss function for the legality classification task and the binary cross-entropy loss function for the violation type identification task.

9. The method for automatically detecting the legality of a privacy policy based on deep learning according to claim 1, characterized in that, Step 2 also includes: during model training, selecting the AdamW optimizer and combining it with a warmup mechanism, while introducing Dropout regularization strategy in both the Convolutional Neural Network (CNN) and the Bidirectional Long Short-Term Memory (BiLSTM) network.

10. A deep learning-based automatic privacy policy legality detection system utilizing the deep learning-based automatic privacy policy legality detection method according to any one of claims 1-9, characterized in that, include: Text Acquisition and Annotation Module: Used to acquire privacy policy texts based on a multi-dimensional corpus screening mechanism driven by industry risk weights, and to construct fine-grained semantic judgment rules based on the determined compliance dimensions and legality status labels, and to generate legality annotation results using a large language model; Model building and training module: used to build a dual-task architecture for legality classification and violation type identification based on a deep learning language model, and to train the model based on the legality annotation dataset obtained in the text acquisition and annotation module; wherein, the dual-task architecture uses the pre-trained language model BERT as the base encoding layer, and uses a convolutional neural network CNN and a bidirectional long short-term memory network BiLSTM for legality classification and violation type identification respectively. Legality identification module: This module is used to input the privacy policy text to be tested into the model trained in the model building and training module, and obtain the legality classification and violation type identification results.

Citation Information

Cited By

  • Application privacy compliance detection method based on large language model and related device

    CN121256817A