A data compliance analysis method
By building a legal knowledge base and compliance questionnaire library, and using natural language processing and machine learning technology to automatically identify and evaluate data compliance, the difficulties in organizing and evaluating legal articles in data compliance analysis are solved, and efficient recommendations and compliance scoring of violation laws are achieved.
Patent Information
- Application Number
- CN202211741455.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-31
AI Technical Summary
There is a lack of unified organization of relevant laws and regulations on data compliance and sorting out key points of violations in existing resources, and the lack of quantitative compliance indicator design, which makes it difficult to evaluate data compliance.
Adopt natural language processing and machine learning technology to build a legal knowledge base, compliance questionnaire database and intelligent suggestions library, automatically identify legal provisions, generate custom compliance questionnaire, analyze data risks and recommend the laws with the highest probability of violations, and conduct form authenticity verification and violation items risk rating.
It realizes automatic recommendation of violation laws and regulations, automatically evaluates the risk of data violations and the degree of compliance, and improves the efficiency and accuracy of data compliance analysis.
Smart Images

Figure CN116401343B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data compliance, and particularly relates to a data compliance analysis method. Background Art
[0002] With the initial improvement of the legal framework for data compliance, the state has put forward clear requirements for data compliance and personal information protection. At the same time, the penalties for data violations at home and abroad have been increased, and the supervision of data security has become more stringent. Therefore, the demand for a comprehensive understanding of data compliance laws is constantly expanding.
[0003] The miscellaneous types of legal materials and the professionalism of the law trouble the ordinary people in learning, understanding, and using the law, making it difficult to turn the law into a weapon to protect their own rights and interests. In the existing resources, there is a lack of unified collation of legal provisions related to data compliance, sorting out of violation key points, and self-recommendation of violation laws. If sorted out, screened, and scored manually, such a large amount of data collation would require a lot of time and manpower and is basically impossible to achieve. In addition, there is no unified standard for data compliance and various dimension evaluation indicators. Therefore, a compliance quantification index design method is needed to solve the above dilemmas. Summary of the Invention
[0004] To solve the problem of the lack of collation, annotation, and splitting of legal provisions related to data compliance in the existing resources, the purpose of the embodiments of the present application is to provide a data compliance analysis method.
[0005] According to the first aspect of the embodiments of the present application, a data compliance analysis method is provided, including:
[0006] Step S101: Set up a legal knowledge base, a compliance questionnaire library, and an intelligent suggestion library;
[0007] Step S102: Select a target law in the legal knowledge base, automatically identify the legal provisions related to data compliance therein, and make annotations;
[0008] Step S103: Select a target module in the compliance questionnaire so that the user can obtain a customized compliance questionnaire form and make selections;
[0009] Step S104: Input a case description, data, and the answered compliance questionnaire into the intelligent suggestion library, analyze the data risks, and recommend several laws with the highest violation probabilities;
[0010] Step S105: Verify the form authenticity of the answered compliance questionnaire, rate the risks of violation items, and score each data dimension to obtain a data compliance score.
[0011] Further, the legal knowledge base includes:
[0012] Domestic legal knowledge base, which contains all valid documents related to data compliance in the country and is classified according to laws, administrative regulations, departmental rules and regulatory documents, as well as national and industry standards;
[0013] International legal knowledge base, which contains important legal documents related to data compliance internationally and is classified by country and region.
[0014] Furthermore, the compliance questionnaire library includes:
[0015] Regulation questionnaire library, which is for each law, splits each article of the law to obtain the compliance questionnaire exclusive to the law;
[0016] Data full-process questionnaire library, which includes a data full-process security assessment questionnaire and a data full-process processing assessment questionnaire. The data full-process security assessment questionnaire is divided into a basic assessment module and a technical ability assessment module. The data full-process processing assessment questionnaire is divided into modules of data collection, data transmission, data storage, data use, data disclosure, data destruction, and entrusted processing according to each data process;
[0017] Among them, the compliance questionnaire library supports customizing questionnaires, selecting corresponding modules from the regulation questionnaire library or the data full-process questionnaire library or adding custom modules, and selecting questions according to preset rules to generate a questionnaire form unique to the project.
[0018] Furthermore, the intelligent advice library includes:
[0019] The first advice library, which collects data, intelligently analyzes the risks contained in the data, and gives several articles with the highest violation probabilities;
[0020] The second advice library, which searches for violation items, obtains the problems corresponding to the violation items and generates a compliance questionnaire, and lists several articles with the highest violation probabilities through the user's selected answers.
[0021] Furthermore, select the target law in the legal knowledge base, automatically identify and mark the legal articles related to data compliance therein, including:
[0022] Perform corpus preprocessing, feature extraction, and classification on the target law to obtain all legal articles related to data compliance;
[0023] Special annotations are made through marking of violation items and splitting of lower-level laws. Among them, the violation items include data security violation items and data processing violation items. The data security violation items include data classification and grading, institutional guarantee, data identification, interface security management, and data leakage prevention. The data processing violation items mainly include data collection, data transmission, data storage, data use, data disclosure, data destruction, and entrusted processing. The splitting of lower-level laws is to automatically search for and extract the ambiguous parts in the existing laws and the violation items involved in them in the lower-level laws through natural language processing and machine learning models.
[0024] Furthermore, the case description is obtained through self-description by the user or description by a third party. The data includes data in text, voice, image, or video format.
[0025] Furthermore, analyze the data risks and recommend several laws with the highest violation probabilities, including:
[0026] Intelligently analyze the risks contained in the data. The intelligent analysis is to extract natural semantic elements from the data through natural language processing, classify the natural semantic elements according to each violation item through a convolutional neural network, mark the violation items with risks in the data, and annotate the natural language elements and violation items corresponding to the data.
[0027] Match the obtained natural semantic elements and violation items with the legal and regulatory elements corresponding to each legal provision to obtain several legal provisions with the highest violation probabilities.
[0028] Furthermore, analyze the data risks and recommend several laws with the highest violation probabilities, including:
[0029] Search for violation items, obtain the problems corresponding to the violation items and generate a compliance questionnaire for selection. The compliance questionnaire automatically searches for the questionnaire template corresponding to the violation item in the questionnaire library based on the violation item input by the user, combines it into the compliance questionnaire for the violation item, and the user makes selections. Conduct semantic element analysis and keyword extraction on the selected compliance questionnaire, classify the natural semantic elements according to each violation item through a convolutional neural network, and mark the natural semantic elements and violation items of the questionnaire selection.
[0030] Match the obtained natural semantic elements and violation items with the legal and regulatory elements corresponding to each legal provision to obtain several legal provisions with the highest violation probabilities.
[0031] Further, the form authenticity verification is to find key inconsistent issues in the custom compliance questionnaire form to detect user input errors or potential fraud behaviors; the risk rating of non-compliant items is to analyze data non-compliant items containing risks based on the form selected by the user, analyze the risk levels of each non-compliant item according to the user's selected answers, and classify the risks; the scores of each data dimension calculate the area of an irregular polygon based on specific non-compliant items, data volume, and data categories, and obtain the scores of each data dimension, where the data dimensions include privacy protection, data security, process standardization, and data confidentiality.
[0032] Further, finding key inconsistent issues in the custom compliance questionnaire form specifically means for the form any question in , calculate the anomaly of the question : :
[0033]
[0034] If , then mark this question as a key inconsistent issue and remind the user to check it carefully.
[0035] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0036] As can be seen from the above embodiments, the present application adopts natural language processing and machine learning technologies, overcomes the problem that users are not clear about the laws and regulations regarding possible violations of their data, and thus achieves the technical effect of automatically recommending violation laws and regulations; overcomes the problem that users do not understand the risk levels of their data, and thus achieves the technical effects of automatic risk rating of data non-compliant items and automatic scoring of data compliance.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0039] Figure 1 is a flowchart of a data compliance analysis method shown according to an exemplary embodiment.
[0040] Figure 2 is a flowchart of step S103 shown according to an exemplary embodiment.
[0041] Figure 3 is a flowchart of form authenticity verification shown according to an exemplary embodiment.
[0042] Figure 4 is a flowchart for obtaining a data compliance score shown according to an exemplary embodiment.
[0043] Figure 5 is a block diagram of a data compliance analysis device shown according to an exemplary embodiment.
[0044] Figure 6 is a schematic diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0045] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0046] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0047] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0048] Figure 1 is a flowchart of a data compliance analysis method shown according to an exemplary embodiment. As Figure 1 shown, the method may include the following steps:
[0049] Step S101: Set up a legal knowledge base, a compliance questionnaire library, and an intelligent suggestion library;
[0050] Step S102: Select a target law in the legal knowledge base, automatically identify the legal provisions related to data compliance therein, and make annotations;
[0051] Step S103: Select a target module in the compliance questionnaire so that the user can obtain a customized compliance questionnaire form and make selections;
[0052] Step S104: Input the case description, data, and the compliant questionnaire after answering into the intelligent suggestion library, analyze the data risks, and recommend several articles with the highest violation probabilities.
[0053] Step S105: Conduct form authenticity verification, risk rating of violation items, and scoring of each data dimension for the compliant questionnaire after answering to obtain a data compliance score.
[0054] As can be seen from the above embodiments, the present application adopts natural language processing and machine learning technologies, overcomes the problem that users are not clear about the articles of law that their data may violate, and thus achieves the technical effect of automatically recommending violation articles of law; overcomes the problem that users do not understand the degree of risk existing in their data, and thus achieves the technical effects of automatic risk rating of data violation items and automatic scoring of data compliance.
[0055] In the specific implementation of step S101, a legal knowledge library, a compliant questionnaire library, and an intelligent suggestion library are set up.
[0056] Specifically, legal articles are stored in the legal knowledge library, including a domestic legal knowledge library and an international legal knowledge library. The domestic legal knowledge library contains all valid documents related to data compliance in the country, and is classified according to laws, administrative regulations, departmental rules and regulatory documents, national and industry standards, etc.; the international legal knowledge library contains important legal documents related to data compliance in the world, and is classified according to countries and regions.
[0057] The compliant questionnaire library includes a regulation questionnaire library and a data full-process questionnaire library. The regulation questionnaire library is for each law, splitting each article of the law to obtain a compliant questionnaire exclusive to the law; the data full-process questionnaire library includes a data full-process security assessment questionnaire and a data full-process processing assessment questionnaire. The data full-process security assessment questionnaire is divided into a basic assessment module and a technical ability assessment module. The data full-process processing assessment questionnaire is divided into modules such as data collection, data transmission, data storage, data use, data disclosure, data destruction, and entrusted processing according to the data process. The compliant questionnaire library supports customizing questionnaires, selecting corresponding modules from the data full-process questionnaire library or adding custom modules, and selecting questions according to preset rules to generate a questionnaire form unique to the project.
[0058] The intelligent suggestion library includes a first suggestion library and a second suggestion library. The first suggestion library collects data, intelligently analyzes the risks contained in the data, and recommends several articles with the highest violation probabilities; the second suggestion library searches for violation items, obtains the corresponding questions for the violation items and generates a compliant questionnaire, and lists several articles with the highest violation probabilities through answering.
[0059] In the specific implementation of step S102, select the target law in the legal knowledge base, automatically identify the legal provisions related to data compliance therein and mark them;
[0060] Specifically, the automatic identification includes corpus preprocessing, feature extraction, and classifier selection of legal provisions to obtain all legal provisions related to data compliance; the special marking includes: violation item marking and lower-level law splitting.
[0061] In one embodiment, the automatic identification and marking are visually displayed, and the relevant legal provisions will be displayed in different colored fonts. When the mouse is moved to the data compliance legal provisions, a floating box will appear, and the floating box includes legal and regulatory elements, violation items, and lower-level law splitting.
[0062] Specifically, the violation items include data security violation items and data processing violation items. The data security violation items include data classification and grading, system guarantee, data identification, interface security management, data leakage prevention, etc. The data processing violation items mainly include data collection, data transmission, data storage, data use, data disclosure, data destruction, entrusted processing, etc.
[0063] Specifically, the lower-level law splitting is automatically searched for and extracted in the lower-level law based on the fuzzy part in the existing legal provisions and the violation items involved through natural language processing and machine learning models.
[0064] In the specific implementation of step S103, select the target module in the compliance questionnaire so that the user can obtain a customized compliance questionnaire form and answer it;
[0065] Specifically, the regulation questionnaire library contains specific questionnaires for all laws related to data compliance. The questionnaires are obtained by splitting each legal provision, analyzing violation items, and setting questions. Among them, splitting according to each legal provision includes splitting according to length and splitting according to content; the data full-process questionnaire library includes data full-process security assessment questionnaires and data full-process processing assessment questionnaires, which are divided into different modules according to each process of data.
[0066] The compliance questionnaire library supports adding customized questionnaires. The user selects specific modules and customized modules in the questionnaire library to generate a customized assessment questionnaire. For example, the user selects the questionnaire related to the data processing link in the data full-process questionnaire library and adds customized questions related to the data processing link to form a customized assessment questionnaire.
[0067] In the specific implementation of step S104, input the case description, data, and the answered compliance questionnaire into the intelligent recommendation library, analyze the data risk and recommend several legal provisions with the highest violation probability;
[0068] Specifically, in the first recommendation library, word segmentation technology and text analysis technology are used to perform semantic analysis on the case description and questionnaire library, match the analysis results with each legal provision, and obtain several legal provisions with the highest violation probability through model training.
[0069] Figure 2 This is the operation process of the intelligent questionnaire library in this application. Refer to Figure 2 As shown, S201 collecting data and S202 intelligently analyzing the risks contained in the data are search method one, and S211 searching for violation items and S212 obtaining the questions corresponding to the implemented violation items and generating a compliance questionnaire for answering are search method two.
[0070] Refer to Figure 2 As shown, S201 collects data, and the data includes case descriptions, data input, compliance questionnaire answering, etc. The questionnaire library supports processing data formats such as text, voice, images, and videos. S202 intelligently analyzes the risks contained in the data. The intelligent analysis extracts natural semantic elements from the input data through natural language processing, classifies the natural semantic elements according to each violation item through a convolutional neural network, marks the violation items with risks in the data, and annotates the natural language elements and violation items corresponding to the data.
[0071] Refer to Figure 2 As shown, S211 searches for violation items, and the violation items need to be selected from the drop-down box of the search bar. The violation items are mainly divided into two categories: data security and data processing; S212 obtains the questions corresponding to the violation items and generates a compliance questionnaire for answering. The compliance questionnaire automatically searches for the questionnaire template corresponding to the violation item in the questionnaire library through the violation item input by the user, combines it into the compliance questionnaire for the violation item, and is answered by the user. Specifically, semantic element analysis and keyword extraction are performed on the answered questionnaire, and the natural semantic elements are classified according to each violation item through a convolutional neural network, and the natural semantic elements and violation items of the questionnaire answer are marked.
[0072] Refer to Figure 2 As shown, S221 is based on several legal provisions with the highest recommended violation probability. By performing semantic analysis and natural semantic element extraction on each legal provision, the legal and regulatory elements corresponding to the legal provision are obtained, and the natural semantic elements and violation items obtained in S202 or S212 are matched with the legal and regulatory elements corresponding to each legal provision, and several (such as ten in one embodiment) legal provisions with the highest violation probability are obtained.
[0073] In the specific implementation of step S105, for the compliance questionnaire after answering, form authenticity verification, violation item risk rating, and scoring of each data dimension are performed to obtain a data compliance score.
[0074] Figure 3 is the flowchart of form authenticity verification in this application. Refer to Figure 3 as shown, S301: Transmit the user's selected answer results to the form authenticity verification system, and the verification system automatically determines whether there are any abnormalities in the filling. The specific steps are as follows:
[0075] First, define any form . Among them, represents the option value corresponding to the th question; ; , is the number of options for the corresponding question.
[0076] Secondly, in order to balance the attribute weights of different questions, the complementary entropy that measures the information uncertainty and ambiguity in the rough set theory is adopted here as a measure of the information gain or uncertainty of categorical data. Its definition is as follows:
[0077]
[0078] Among them, is the number of options for question , is the complement of , ; represents the probability of the equivalence class of within the universe ; represents the probability of the complement of appearing in the universe .
[0079] Specifically, the abnormality weight of any question r can be defined as follows:
[0080]
[0081] Among them, represents the weighted weight of question in the abnormality measure.
[0082] For any form to be tested, use similarity to find its neighborhood data , that is, is the set of neighborhood data, satisfying ; represents the similarity with the th nearest form. The distance similarity formula between forms can be expressed as:
[0083]
[0084] Among them, is the option value of the th question of the form; ⊕ represents exclusive OR; x represents dot product.
[0085] Furthermore, calculate the local anomaly degree of the sample to be detected , and its formula can be expressed as:
[0086]
[0087] If , then there is an anomaly in the form and it needs to be checked and re-authenticated, otherwise it is not necessary.
[0088] S311: If an anomaly occurs during the verification process, highlight the key inconsistent issues in red and remind the user to check the filled content.
[0089] The steps to find the key inconsistent issues are as follows:
[0090] For any question in this form , it is necessary to calculate the anomaly degree of the question , which can be expressed as:
[0091]
[0092] If , then mark this question as a key inconsistent issue and remind the user to check it carefully.
[0093] S312: After verification, the user signs and submits the form authenticity commitment agreement again.
[0094] The form authenticity commitment agreement includes the commitment to follow the principle of integrity when filling out the form, the content marked for key inspection, and the behavior record of the user changing options when verifying key inspection questions.
[0095] S321: If no anomaly appears in the verification result, it is handed over to the scoring system for compliance scoring. Jump to Figure 4 the compliance scoring process shown.
[0096] In the case where no anomaly appears in the verification result , then there is no anomaly in this form.
[0097] Refer to Figure 4 shown, S401: Analyze the data violation items of the risks contained in the form selected by the user and classify the violation items according to the risk level.
[0098] The user-selected form is the custom compliance questionnaire form generated in step S103, which will not be elaborated here.
[0099] The data violation items with risks are obtained by analyzing each question in the form. Specifically, by extracting natural semantic elements from each question in the questionnaire library, each question corresponds to one or more data violation items. Each question has a corresponding weight, and each answer to the question has a corresponding score.
[0100] Specifically, after the user finishes answering, the system calculates the risk scores of each violation item based on the weights of each question and the scores of the answers. When the risk score is higher than a certain score, the violation item is marked as a risk-containing violation item, and according to the benchmark, the risk-containing violation items are classified into severe, high, medium, and low levels.
[0101] S402: According to the violation items involved in each data dimension, calculate the area of an irregular polygon to obtain the score of each data dimension.
[0102] The data dimensions include privacy protection, data security, process standardization, and data confidentiality.
[0103] Specifically, each data dimension involves different data violation items. A radar chart is formed based on specific violation items and variables such as data volume and data category, and the score of a certain data dimension is obtained by calculating the area of the radar chart.
[0104] S403: According to the scores of each data dimension, calculate the area of an irregular polygon to obtain the data compliance score.
[0105] The data compliance is obtained by calculating the area of a radar chart composed of four variables: privacy protection, data security, process standardization, and data confidentiality. The data compliance score is a real number between 0 and 100.
[0106] Finally, the data compliance quantification index of a system can enable users to intuitively understand the compliance degree of their data. In addition, by sorting out relevant laws and regulations on data compliance, sorting out violation items, and recommending violation laws, it can help users learn, understand, and use the law, and understand several laws with the highest probability of their data violations.
[0107] Corresponding to the foregoing embodiments of the data compliance analysis method, the present application also provides an embodiment of a data compliance analysis device.
[0108] Figure 2 is a block diagram of a data compliance analysis device shown according to an exemplary embodiment. Refer to Figure 2 and the device may include:
[0109] A setting module 21 for setting up a legal knowledge base, a compliance questionnaire library, and an intelligent advice library;
[0110] An identification module 22 for selecting target laws in the legal knowledge base, automatically identifying legal provisions related to data compliance therein, and making annotations;
[0111] A selection module 23 for selecting a target module in the compliance questionnaire, so that the user can obtain a customized compliance questionnaire form and make selections;
[0112] An analysis module 24 for inputting case descriptions, data, and the answered compliance questionnaire into the intelligent advice library, analyzing data risks, and recommending several laws with the highest violation probabilities;
[0113] A scoring module 25 for performing form authenticity verification, violation item risk rating, and scoring of each data dimension on the answered compliance questionnaire to obtain a data compliance score.
[0114] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0115] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0116] Correspondingly, the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the data compliance analysis method as described above. As Figure 4 shown, it is a hardware structure diagram of a data compliance analysis system provided by an embodiment of the present invention in any device with data processing capabilities. Except for Figure 4 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated herein.
[0117] Correspondingly, the present application also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the data compliance analysis method as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0118] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0119] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A data compliance analysis method, characterized in that, Including: Step S101: Set up a legal knowledge base, a compliance questionnaire library, and an intelligent advice library; Step S102: Select a target law in the legal knowledge base, automatically identify the legal provisions related to data compliance therein, and mark them; Step S103: Select a target module in the compliance questionnaire to enable the user to obtain a customized compliance questionnaire form and make selections; Step S104: Input a case description, data, and the answered compliance questionnaire into the intelligent advice library, analyze the data risks, and recommend several legal articles with the highest violation probabilities; Step S105: Conduct form authenticity verification, risk rating of violation items, and scoring of each data dimension for the answered compliance questionnaire to obtain a data compliance score; Among them, the form authenticity verification is to search for key inconsistent problems in the customized compliance questionnaire form to detect user input errors or potential fraud behaviors; the risk rating of violation items is to analyze the data violation items containing risks based on the user's self-selected form, analyze the risk levels of each violation item according to the user's selections, and classify the risks; The scoring of each data dimension calculates the area of an irregular polygon based on specific violation items, data volume, and data categories, and obtains the scores of each data dimension, where the data dimensions include privacy protection, data security, process standardization, and data confidentiality; Find key inconsistency issues in the custom compliance questionnaire form, specifically for the form for any question , calculate the anomaly degree of the question : : , Among them is the form to be inspected 's neighborhood data is the option value of the th question of the form If , then mark this problem as a key inconsistency problem and remind the user to check it carefully.
2. The method according to claim 1, wherein The legal knowledge base includes: A domestic legal knowledge base, which contains all effective documents related to data compliance in the country and is classified according to laws, administrative regulations, departmental rules and regulatory documents, and national and industry standards; An international legal knowledge base, which contains important legal documents related to data compliance internationally and is classified by country and region.
3. The method according to claim 1, wherein The compliance questionnaire library includes: A regulation questionnaire library, which splits the legal provisions of each law to obtain a compliance questionnaire exclusive to the law; A data full-process questionnaire library, which includes a data full-process security assessment questionnaire and a data full-process processing assessment questionnaire. The data full-process security assessment questionnaire is divided into a basic assessment module and a technical ability assessment module. The data full-process processing assessment questionnaire is divided into modules for data collection, data transmission, data storage, data use, data disclosure, data destruction, and entrusted processing according to each data process; Among them, the compliance questionnaire library supports customized questionnaires. Select corresponding modules from the data full-process questionnaire library or add customized modules, and select questions according to preset rules to generate a questionnaire form unique to the project.
4. The method according to claim 1, wherein The intelligent advice library includes: A first advice library, which collects data, intelligently analyzes the risks contained in the data, and gives several legal articles with the highest violation probabilities; A second advice library, which searches for violation items, obtains the corresponding questions for the violation items, generates a compliance questionnaire, and lists several legal articles with the highest violation probabilities through the user's selections.
5. The method according to claim 1, wherein Selecting a target law in the legal knowledge base, automatically identifying the legal provisions related to data compliance therein, and marking them includes: Preprocess the corpus, extract features, and classify the target law to obtain all legal provisions related to data compliance; Perform special annotation through violation item annotation and lower-level law splitting. Among them, the violation items include data security violation items and data processing violation items. The data security violation items include data classification and grading, system guarantee, data identification, interface security management, and data leakage prevention. The data processing violation items mainly include data collection, data transmission, data storage, data use, data disclosure, data destruction, and entrusted processing. The lower-level law splitting is to automatically search for and extract the ambiguous parts in the existing legal provisions and the violation items involved in them in the lower-level laws through natural language processing and machine learning models.
6. The method according to claim 1, wherein The case description shown is obtained through self-description by the user or third-party description. The data includes data in text, voice, image, or video format.
7. The method according to claim 1, characterized in that Analyze the data risks and recommend several legal provisions with the highest violation probabilities, including: Intelligently analyze the risks contained in the data. The intelligent analysis is to extract natural semantic elements from the data through natural language processing, classify the natural semantic elements according to each violation item through a convolutional neural network, mark the violation items with risks in the data, and annotate the natural language elements and violation items corresponding to the data; Match the obtained natural semantic elements and violation items with the legal and regulatory elements corresponding to each legal provision to obtain several legal provisions with the highest violation probabilities.
8. The method according to claim 1, characterized in that, Analyze the data risks and recommend several legal provisions with the highest violation probabilities, including: Search for violation items, obtain the problems corresponding to the violation items, and generate a compliance questionnaire for selection. The compliance questionnaire is to automatically search for the questionnaire template corresponding to the violation item in the questionnaire library based on the violation item input by the user, combine it into the compliance questionnaire for the violation item, and let the user make selections. Perform semantic element analysis and keyword extraction on the compliance questionnaire after selection, classify the natural semantic elements according to each violation item through a convolutional neural network, and mark the natural semantic elements and violation items of the questionnaire selection; Match the obtained natural semantic elements and violation items with the legal and regulatory elements corresponding to each legal provision to obtain several legal provisions with the highest violation probabilities.
Citation Information
Patent Citations
Method and device for determining abnormal data
CN112597255A
Data compliance self-inspection method and device
CN114416958A