A data element automatic compliance verification method, device, equipment and medium
Patent Information
- Application Number
- CN202610729759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]现有数据合规校验技术多采用硬编码方式配置,灵活性差,难以适配多行业合规需求,法规与政策更新时响应慢、维护成本高,并且过度依赖人工审核与规则制定,整体效率低、人力成本高,且易出现漏判误判的情况,难以满足数据要素全流程自动化合规校验的需求
[0015] Compared with existing technologies, the automated compliance verification method for data elements provided in this invention has the following advantages: It constructs an industry compliance graph based on multi-source compliance knowledge; wherein the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; it models the entire lifecycle of the data elements to be verified and extracts metadata for each processing stage; it matches target compliance rules applicable to the metadata from the industry compliance graph; based on the target compliance rules, it performs automated compliance verification on the data elements and outputs the verification results. This invention can effectively improve the efficiency, accuracy, and real-time performance of compliance verification.
Smart Images

Figure CN122736383A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data compliance verification technology, and in particular to an automated compliance verification method, apparatus, device, and medium for data elements. Background Technology
[0002] With the continuous development of the digital economy, data has become a core factor of production. The construction, circulation, trading, and assetization of data element markets are rapidly advancing, but the resulting data compliance issues are becoming increasingly serious. Throughout the entire lifecycle of data elements—from collection and storage to processing, circulation, and application—compliance must be maintained with laws and regulations such as the Personal Information Protection Law, the Data Security Law, the Cybersecurity Law, and the Measures for Security Assessment of Cross-border Data Export, as well as specific compliance standards for different industries such as finance, healthcare, and telecommunications. This places higher demands on the comprehensiveness, accuracy, and timeliness of compliance management.
[0003] Existing data compliance verification technologies mostly adopt hard-coded configuration, which is inflexible, difficult to adapt to the compliance needs of multiple industries, slow to respond to updates of laws and policies, and has high maintenance costs. In addition, they rely too much on manual review and rule making, resulting in low overall efficiency, high labor costs, and a tendency to miss or misjudge cases, making it difficult to meet the needs of automated compliance verification of data elements throughout the entire process. Summary of the Invention
[0004] This invention provides an automated compliance verification method for data elements. By constructing a multi-dimensional knowledge graph, it can effectively improve the efficiency, accuracy, and real-time performance of compliance verification.
[0005] In a first aspect, embodiments of the present invention provide an automated compliance verification method for data elements, comprising: An industry compliance graph is constructed based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; The data elements to be verified are modeled throughout their entire lifecycle, and metadata is extracted from each processing step. Match the target compliance rules applicable to the metadata from the industry compliance map; Based on the target compliance rules, automated compliance verification is performed on the data elements, and the verification results are output.
[0006] Furthermore, the construction of the industry compliance graph based on multi-source compliance knowledge includes: Compliance text data is collected from multiple channels and preprocessed. A large language model is used to extract rule elements from the preprocessed compliant text data; Identify rule entities from the extracted rule elements and determine the relationships between the rule entities; By using the rule entities as nodes and the relationships as edges, an industry compliance graph is constructed.
[0007] Furthermore, the data elements to be verified undergo full lifecycle process modeling, and metadata for each processing stage is extracted, including: The data element lifecycle is divided into five stages: data collection, data storage, data processing, data circulation, and data application. Metadata management tools are used to extract metadata for each stage of the lifecycle; Based on the extracted metadata, a data lineage relationship is constructed to obtain the complete lineage link of data from the source to the application.
[0008] Furthermore, the step of matching target compliance rules applicable to the metadata from the industry compliance graph includes: For the metadata of each step, based on the data type, processing scenario and responsible entity of the metadata, calculate the semantic similarity with each compliance rule; The compliance rules whose semantic similarity exceeds a preset threshold are used as the initial compliance rules; Based on the node relationships in the industry compliance graph, a graph neural network is used to perform deep reasoning on the initial compliance rules to identify rule conflicts, rule redundancy, and hidden risks, thereby obtaining the final target compliance rules.
[0009] Furthermore, the method also includes: A streaming computing framework is used to collect data flow events in real time; these flow events include, but are not limited to, data access events, data modification events, data transmission events, and data sharing events. Perform real-time compliance checks on each process flow event and calculate a risk score; Based on the risk score and the preset multi-level warning thresholds, the corresponding warning mechanism is triggered.
[0010] Furthermore, the method also includes: Construct a compliance intelligent question-answering system based on RAG retrieval enhancement, and use the industry compliance graph as the knowledge base of the compliance intelligent question-answering system; When a compliance inquiry is received from a user, the system retrieves information from the knowledge base, inputs the retrieved knowledge into the large language model, and outputs the corresponding answer to the compliance question.
[0011] Furthermore, the method also includes: A risk prediction model is constructed and trained based on time series forecasting algorithms and machine learning classification algorithms. The current compliance risk score, the frequency of historical violations, the trend of data sensitivity changes, the updates of laws and regulations, and the changes in business scenarios are used as input features and input into the risk prediction model. The model outputs the risk probability and risk level predicted by the model for a future preset period. Proactive risk warnings are issued based on the risk probability and the risk level.
[0012] Secondly, embodiments of the present invention provide an automated compliance verification device for data elements, comprising: The compliance graph construction module is used to construct an industry compliance graph based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; The full-process modeling module is used to perform full lifecycle process modeling on the data elements to be verified and extract metadata for each processing step. The compliance rule extraction module is used to match target compliance rules applicable to the metadata from the industry compliance map. The compliance verification module is used to perform automated compliance verification on data elements based on the target compliance rules and output the verification results.
[0013] Thirdly, embodiments of the present invention provide an electronic device, comprising: Memory, used to store computer programs; A processor for executing the computer program; Wherein, when the processor executes the computer program, it implements the data element automated compliance verification method described in any of the first aspects above.
[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed, implements the automated compliance verification method for data elements as described in any of the first aspects above.
[0015] Compared with existing technologies, the automated compliance verification method for data elements provided in this invention has the following advantages: It constructs an industry compliance graph based on multi-source compliance knowledge; wherein the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; it models the entire lifecycle of the data elements to be verified and extracts metadata for each processing stage; it matches target compliance rules applicable to the metadata from the industry compliance graph; based on the target compliance rules, it performs automated compliance verification on the data elements and outputs the verification results. This invention can effectively improve the efficiency, accuracy, and real-time performance of compliance verification. Attached Figure Description
[0016] To more clearly illustrate the technical features of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating an automated compliance verification method for data elements provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an automated compliance verification device for data elements provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0021] In a first aspect, embodiments of the present invention provide an automated compliance verification method for data elements, see [link to relevant documentation]. Figure 1 This is a flowchart illustrating an embodiment of an automated compliance verification method for data elements provided by the present invention.
[0022] like Figure 1 As shown, the method includes the following steps: S1: Construct an industry compliance graph based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; S2: Perform full lifecycle process modeling on the data elements to be verified and extract metadata for each processing step; S3: Match the target compliance rules applicable to the metadata from the industry compliance map; S4: Based on the target compliance rules, perform automated compliance verification on the data elements and output the verification results.
[0023] Specifically, the process begins by collecting multi-source compliance data from multiple channels, extracting various compliance requirements from this data, abstracting them into rule entities, establishing relationships between these rule entities, and finally forming a structured and reasonable industry compliance graph with rule entities as nodes and relationships between them as edges. This provides unified and authoritative knowledge support for subsequent compliance verification.
[0024] Furthermore, the system performs full lifecycle process modeling on the data elements to be verified, dividing the actual processing of data elements in the business system into several standard processing steps to form a unified and supervised lifecycle framework. On this basis, for each processing step in the full lifecycle, the system automatically collects, parses, and structures the data to extract compliance-related metadata. The extracted metadata is used to comprehensively characterize the data elements' own attributes, processing methods, flow paths, access control, and usage intent, forming standardized feature information that can be used for compliance matching, risk identification, and graph reasoning.
[0025] Furthermore, using the data attributes, processing scenarios, responsible entities, and flow behaviors represented by metadata as matching criteria, the rule entities and their relationships are retrieved, filtered, and deeply verified in the industry compliance map. The target compliance rules that are compatible with the current data elements and processing links are selected, providing directly executable rule basis for subsequent automated compliance verification.
[0026] Finally, based on the matched target compliance rules, automated compliance verification is performed on the data elements, and the verification results are output. The verification results are divided into three levels, including full compliance (i.e. all applicable rules are met and there are no violations), partial compliance (i.e. some rules are met, but there are some violations or risks), and non-compliance (i.e. there are serious violations that require immediate rectification).
[0027] It should be noted that for some compliant and non-compliant cases, a corresponding compliance risk score needs to be calculated and output along with the verification results. The risk score calculation formula is as follows: ; in, This is the risk score. The severity of the violation is assessed based on legal provisions and precedents. To influence the weighting, factors such as data sensitivity level and user scale are considered. This refers to the compliance gap, i.e., the degree of non-compliance.
[0028] In summary, this invention constructs an industry compliance graph, structuring scattered and unstructured compliance rules into a scalable knowledge graph. This significantly improves the management efficiency and update flexibility of compliance rules, further accurately matches target compliance rules, and completes automated compliance verification based on these target rules. Compared to traditional compliance checks that rely on manual review and static hard rule matching, this invention effectively addresses the technical shortcomings of existing compliance verification methods, such as fragmented rules, one-sided matching, high manual costs, and low verification efficiency. It significantly improves the accuracy of compliance rule adaptation and the coverage of risk identification, achieving full-process, automated compliance judgment of data elements, reducing review errors caused by human intervention, and improving the real-time performance, reliability, and universality of compliance verification. This provides reliable technical support for the standardized and secure flow of data elements within the industry.
[0029] In one optional implementation, the construction of the industry compliance graph based on multi-source compliance knowledge includes: Compliance text data is collected from multiple channels and preprocessed. A large language model is used to extract rule elements from the preprocessed compliant text data; Identify rule entities from the extracted rule elements and determine the relationships between the rule entities; By using the rule entities as nodes and the relationships as edges, an industry compliance graph is constructed.
[0030] Specifically, compliant text data is first collected from various channels, including legal and regulatory databases, industry standard libraries, internal corporate documents, regulatory announcements, and court precedents. The collected text data undergoes preprocessing operations such as cleaning, noise reduction, word segmentation, and format standardization to remove invalid information and unify the text structure, providing a high-quality data foundation for subsequent rule extraction and entity recognition.
[0031] Large language models, such as GPT-4 and Claude, are used to automatically extract rules from preprocessed compliant text, accurately extracting the key elements of the rules from the text. For example, the extracted key elements are:
[0032] ; in, The subject of the rules (who must comply). For the object of the rule (what data or business it applies to), This refers to the rule conditions (under what circumstances they apply). Rules are defined as actions that must be performed or are prohibited from being performed. Consequences of violating the rules (the consequences of breaking the rules). The source of the rules (the origin and basis of the rules). Effective date (the date when the rule officially takes effect). The scope of jurisdiction (the geographical area, industry, regulatory field, or subject to which the rules apply).
[0033] Based on named entity recognition and relationship extraction technology, key rule entities such as law name, clause number, data type, business scenario, responsible entity, and industry type are further identified from the extracted rules. The relationships between entities are determined, including but not limited to hierarchical relationships (e.g., law > law > rule > normative document), citation relationships (e.g., law A cites law B), conflict relationships (e.g., law A conflicts with law B), supplementary relationships (e.g., law A supplements law B), and implementation relationships (e.g., law A requires the implementation of a certain measure), forming a complete entity and relationship system.
[0034] The rule entities identified above are used as nodes in the knowledge graph, and the relationships between entities are used as edges in the graph. These are then imported into a graph database for storage and management, thereby constructing an industry compliance graph that covers multiple industries, multiple dimensions, and can be dynamically expanded. For example, node types include, but are not limited to, legal nodes, regulatory nodes, rule nodes, data type nodes, business scenario nodes, responsible entity nodes, industry nodes, and penalty case nodes. Edge types include, but are not limited to, reference edges, conflict edges, supplementary edges, implementation edges, applicable edges, prohibition edges, and penalty edges.
[0035] It should be noted that when rules need to be updated, it is only necessary to add, modify or delete the corresponding nodes and relationships in the graph, which can greatly shorten the update cycle and improve the flexibility of rule maintenance.
[0036] Optionally, after the industry compliance graph is constructed, a graph quality assessment model can be established to evaluate the quality of the industry compliance graph from dimensions such as coverage, accuracy, consistency, and timeliness, and to manually review and correct low-quality nodes and edges.
[0037] This embodiment constructs an industry compliance graph, which structures scattered and unstructured compliance rules into a unified, reasonable, and scalable knowledge graph. It can simultaneously support multiple industries such as finance, healthcare, government affairs, and e-commerce, significantly improving the management efficiency and update flexibility of compliance rules.
[0038] In one optional implementation, the data elements to be verified undergo full lifecycle process modeling, and metadata for each processing stage is extracted, including: The data element lifecycle is divided into five stages: data collection, data storage, data processing, data circulation, and data application. Metadata management tools are used to extract metadata for each stage of the lifecycle; Based on the extracted metadata, a data lineage relationship is constructed to obtain the complete lineage link of data from the source to the application.
[0039] Specifically, the flow process of data elements is standardized and divided into five core stages: data acquisition (obtaining raw data from external data sources), data storage (storing data in data lakes, data warehouses, and other storage systems), data processing (cleaning, transforming, and aggregating data), data circulation (transferring, sharing, and trading data between systems), and data application (using data for business analysis, model training, decision support, etc.). This forms a standardized process model covering the entire process of data from generation to use.
[0040] Metadata management tools (such as Apache Atlas, DataHub, etc.) are used to automatically extract metadata from the five lifecycle stages mentioned above, including data source information (data source system, collection time, collection method), data attribute information (data type, sensitivity level, data volume, update frequency), data processing information (processing method, processing algorithm, processing parameters, processing results), data circulation information (transmission protocol, transmission path, recipient, authorization scope), and data application information (application scenario, accessing users, access permissions, purpose of use). This achieves automated, comprehensive, and standardized collection of metadata, ensuring that the status and behavior of data elements at each stage are quantifiable and traceable.
[0041] Next, based on the extracted full-link metadata, a data lineage is constructed, identifying and recording the upstream source (the source of the current data and the pre-processing steps), downstream destination (which downstream systems or applications use the current data), and complete flow path of each data node, forming a complete lineage from the data source to the final application.
[0042] This embodiment achieves visualization, traceability, and quantification of the data element flow process through full lifecycle modeling, automated metadata extraction, and data lineage construction, providing an accurate and complete data context for subsequent compliance rule matching, real-time monitoring, and risk reasoning.
[0043] In one optional implementation, matching target compliance rules applicable to the metadata from the industry compliance graph includes: For the metadata of each step, based on the data type, processing scenario and responsible entity of the metadata, calculate the semantic similarity with each compliance rule; The compliance rules whose semantic similarity exceeds a preset threshold are used as the initial compliance rules; Based on the node relationships in the industry compliance graph, a graph neural network is used to perform deep reasoning on the initial compliance rules to identify rule conflicts, rule redundancy, and hidden risks, thereby obtaining the final target compliance rules.
[0044] Specifically, for the metadata corresponding to each stage of the data element's lifecycle, the data type, processing scenario, and responsible entity in the metadata are used as core matching dimensions. Semantic similarity is calculated between these dimensions and the rule objects, rule conditions, and rule subjects in the industry compliance map. A weighted fusion is then used to obtain the overall matching score between the data and the rule, achieving preliminary rule screening. The calculation formula is as follows: ; in, For data With rules semantic similarity, For data types, As the object of the rule, To handle the scene, As a rule condition, As the responsible party, As the subject of the rules, For the weighting coefficients, satisfying + β + γ =1.
[0045] The calculated semantic similarity is compared with a preset threshold, and compliance rules that exceed the threshold are filtered out as initial compliance rules applicable to the current data elements to be verified. The preset threshold can be flexibly set and adjusted according to the business scenario and actual verification accuracy requirements. The specific value of the threshold is not limited here.
[0046] Based on the relationships between nodes in the industry compliance graph, graph neural networks (such as GraphSAGE, GAT, etc.) are used to perform deep reasoning and correlation analysis on the initial compliance rules. By aggregating the information of neighboring nodes, the embedding vector of each node is calculated. The calculation formula is as follows: ; in, Let v be the embedding vector of node v at layer k+1. Let v be the embedding vector of node v at the k-th layer. Let v be the set of neighboring nodes. Let u be the vector of the neighbor node at the kth layer. For aggregate functions, For learnable parameters, This is the activation function.
[0047] By aggregating the node's own information and neighboring node information using the above formula, the node embedding vector can be calculated, enabling multi-hop association analysis to identify implicit compliance risks, rule conflicts, and rule redundancy, thereby obtaining the final target compliance rule and completing complex compliance judgments that cannot be achieved by traditional single rule matching.
[0048] This embodiment combines preliminary semantic similarity matching with deep reasoning via graph neural networks, which not only ensures the efficiency and coverage of rule retrieval, but also effectively eliminates rule conflicts and redundancy, uncovers hidden risks that traditional matching methods cannot identify, and significantly improves the accuracy of compliance rules.
[0049] In an optional implementation, the method further includes: A streaming computing framework is used to collect data flow events in real time; these flow events include, but are not limited to, data access events, data modification events, data transmission events, and data sharing events. Perform real-time compliance checks on each process flow event and calculate a risk score; Based on the risk score and the preset multi-level warning thresholds, the corresponding warning mechanism is triggered.
[0050] Specifically, streaming computing frameworks, such as Apache Flink and Apache Spark Streaming, are used to collect and perceive the flow behavior of data elements throughout their entire lifecycle in real time. The monitored data flow events include, but are not limited to, data access events, data modification events, data transmission events, and data sharing events. Among them, data access events are user-initiated access operations on data, data modification events are data modification, deletion, and other editing operations, data transmission events are data transmission operations between different systems, and data sharing events are data sharing operations provided to external entities.
[0051] For each collected flow event, the system immediately performs a real-time compliance check with a millisecond response. First, it quickly retrieves the compliance rules that match the current flow event from the industry compliance map. Then, it compares and verifies the event behavior with the rule requirements. At the same time, it calculates the corresponding risk score based on the severity of the violation, the impact weight, and the compliance gap, providing a quantitative basis for subsequent risk classification and handling.
[0052] To ensure real-time performance, this embodiment further employs rule indexing (creating an inverted index for rules to accelerate rule retrieval), graph caching (caching frequently accessed graph nodes and paths), and parallel computing (parallelizing rule matching and reasoning) to optimize the performance of the rule matching and graph reasoning process, ensuring that compliance checks are completed within the specified time.
[0053] After risk scoring is completed, events are classified and handled according to preset multi-level warning thresholds, forming a multi-level warning mechanism: when the risk score exceeds the high threshold, it is judged as a level 1 high-risk warning, and the system directly and automatically blocks the current operation; when the risk score is between the medium and high thresholds, it is judged as a level 2 medium-risk warning, and the system retains the event record and generates a risk report; when the risk score is between the low and medium thresholds, it is judged as a level 3 low-risk warning, and the system only records the event for subsequent analysis.
[0054] The multi-level early warning thresholds can be flexibly set and adjusted according to business scenarios and actual needs, and no specific value of the thresholds is limited here.
[0055] For events that trigger a Level 1 high-risk warning, the system automatically blocks operations while generating a detailed report containing event information, evidence of violation, and risk level, and pushes it to the compliance team for review. The compliance team can confirm the violation and maintain the blocking status, initiate the rectification process, or determine it as a false alarm and lift the blocking, while updating the compliance graph or rule base accordingly, forming a closed-loop management of monitoring, early warning, handling, and optimization.
[0056] This embodiment achieves real-time monitoring of the entire data flow process through streaming computing. Combined with compliance checks and multi-level early warning mechanisms, it can promptly detect and intercept compliance risks during the dynamic flow of data elements, thereby improving the automation level and efficiency of compliance monitoring.
[0057] In an optional implementation, the method further includes: Construct a compliance intelligent question-answering system based on RAG retrieval enhancement, and use the industry compliance graph as the knowledge base of the compliance intelligent question-answering system; When a compliance inquiry is received from a user, the system retrieves information from the knowledge base, inputs the retrieved knowledge into the large language model, and outputs the corresponding answer to the compliance question.
[0058] Specifically, to improve the intelligence level and response efficiency of compliance management, this embodiment constructs a compliance intelligent question-and-answer system based on the RAG (Retrieval-Augmented Generation) retrieval enhancement generation architecture. The established industry compliance graph is used as the core knowledge base of the system, storing knowledge such as laws and regulations, industry standards, compliance rules, regulatory cases, and clause interpretations in the form of a structured graph, providing an accurate and traceable knowledge source for the question-and-answer process.
[0059] Users can submit compliance questions to the compliance intelligent question-and-answer system. When the system receives a compliance inquiry from a user, it first performs a question understanding step, semantically parsing and identifying the user's intent to determine the question type, such as compliance consultation, risk assessment, rule query, clause interpretation, or rectification suggestions. Subsequently, based on the understood question intent, the system performs precise knowledge retrieval in the industry compliance graph, retrieving relevant compliance clauses, applicable rules, related cases, and official explanations. After acquiring the relevant knowledge, the retrieved compliance knowledge is input into a large language model, logically integrated and organized in conjunction with the question context, generating a clear, well-founded, and regulatory-compliant structured answer. At the same time, the system verifies the accuracy and compliance of the generated answer, checking whether the answer content is consistent with authoritative knowledge in the industry compliance graph to ensure that the output result is compliant and reliable.
[0060] Finally, the system provides the verified answer back to the user, achieving compliant intelligent interaction and decision support in the form of natural language.
[0061] This embodiment combines industry compliance mapping with RAG technology and large language models to achieve rapid retrieval, intelligent integration, and reliable output of compliance knowledge, enabling efficient reuse and unified management of compliance knowledge.
[0062] In an optional implementation, the method further includes: A risk prediction model is constructed and trained based on time series forecasting algorithms and machine learning classification algorithms. The current compliance risk score, the frequency of historical violations, the trend of data sensitivity changes, the updates of laws and regulations, and the changes in business scenarios are used as input features and input into the risk prediction model. The model outputs the risk probability and risk level predicted by the model for a future preset period. Proactive risk warnings are issued based on the risk probability and the risk level.
[0063] Specifically, a compliance risk prediction model is constructed based on time series prediction algorithms and machine learning classification algorithms. The time series prediction algorithms can be ARIMA, LSTM, etc., and the machine learning classification algorithms can be random forest, XGBoost, etc., so that the model has both time series trend fitting and risk classification judgment capabilities.
[0064] During the model's use, multidimensional features are used as inputs to the risk prediction model, including current compliance risk scores, frequency of historical violations, trends in data sensitivity, updates to laws and regulations, and changes in business scenarios, comprehensively covering influencing factors such as data status, historical behavior, external rules, and business changes.
[0065] The above input features are fed into the trained risk prediction model. The model infers and calculates compliance risks for a preset future period (such as the next 7 days or the next 30 days), and outputs the corresponding risk probability and risk level, which are used to quantitatively characterize the likelihood and severity of compliance risks occurring in the future period.
[0066] The system performs proactive risk warnings based on the risk probability and risk level output by the model, issuing early warnings for high-probability and high-level potential risks so that compliance managers can take preventive measures in advance to avoid or reduce the probability of compliance risks occurring.
[0067] This embodiment achieves advanced prediction of compliance risks by integrating time series analysis and machine learning classification. It can accurately identify potential future risks based on multi-dimensional features, effectively improving the foresight and initiative of risk prevention and control.
[0068] Secondly, embodiments of the present invention provide an automated compliance verification device for data elements, see [link to relevant documentation]. Figure 2 This is a schematic diagram of an embodiment of an automated compliance verification device for data elements provided by the present invention.
[0069] like Figure 2 As shown, the device includes: The compliance graph construction module 21 is used to construct an industry compliance graph based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; The metadata management module 22 is used to perform full lifecycle process modeling on the data elements to be verified and extract metadata for each processing step. The compliance rule matching module 23 is used to match target compliance rules applicable to the metadata from the industry compliance map; The compliance verification module 24 is used to perform automated compliance verification on data elements based on the target compliance rules and output the verification results.
[0070] In one optional implementation, the construction of the industry compliance graph based on multi-source compliance knowledge includes: Compliance text data is collected from multiple channels and preprocessed. A large language model is used to extract rule elements from the preprocessed compliant text data; Identify rule entities from the extracted rule elements and determine the relationships between the rule entities; By using the rule entities as nodes and the relationships as edges, an industry compliance graph is constructed.
[0071] In one optional implementation, the data elements to be verified undergo full lifecycle process modeling, and metadata for each processing stage is extracted, including: The data element lifecycle is divided into five stages: data collection, data storage, data processing, data circulation, and data application. Metadata management tools are used to extract metadata for each stage of the lifecycle; Based on the extracted metadata, a data lineage relationship is constructed to obtain the complete lineage link of data from the source to the application.
[0072] In one optional implementation, matching target compliance rules applicable to the metadata from the industry compliance graph includes: For the metadata of each step, based on the data type, processing scenario and responsible entity of the metadata, calculate the semantic similarity with each compliance rule; The compliance rules whose semantic similarity exceeds a preset threshold are used as the initial compliance rules; Based on the node relationships in the industry compliance graph, a graph neural network is used to perform deep reasoning on the initial compliance rules to identify rule conflicts, rule redundancy, and hidden risks, thereby obtaining the final target compliance rules.
[0073] In an optional embodiment, the device is further configured to: A streaming computing framework is used to collect data flow events in real time; these flow events include, but are not limited to, data access events, data modification events, data transmission events, and data sharing events. Perform real-time compliance checks on each process flow event and calculate a risk score; Based on the risk score and the preset multi-level warning thresholds, the corresponding warning mechanism is triggered.
[0074] In an optional embodiment, the device is further configured to: Construct a compliance intelligent question-answering system based on RAG retrieval enhancement, and use the industry compliance graph as the knowledge base of the compliance intelligent question-answering system; When a compliance inquiry is received from a user, the system retrieves information from the knowledge base, inputs the retrieved knowledge into the large language model, and outputs the corresponding answer to the compliance question.
[0075] In an optional embodiment, the device is further configured to: A risk prediction model is constructed and trained based on time series forecasting algorithms and machine learning classification algorithms. The current compliance risk score, the frequency of historical violations, the trend of data sensitivity changes, the updates of laws and regulations, and the changes in business scenarios are used as input features and input into the risk prediction model. The model outputs the risk probability and risk level predicted by the model for a future preset period. Proactive risk warnings are issued based on the risk probability and the risk level.
[0076] It should be noted that the automated compliance verification device for data elements provided in this embodiment of the invention is used to execute all the process steps of the automated compliance verification method for data elements in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0077] Thirdly, embodiments of the present invention provide an electronic device, see [link to previous document]. Figure 3 The diagram shown is a structural schematic of an electronic device provided in an embodiment of the present invention.
[0078] like Figure 3 As shown, the device includes: Memory 31 is used to store computer programs; Processor 32 is used to execute the computer program; When the processor 32 executes the computer program, it implements the data element automated compliance verification method as described in any of the above embodiments.
[0079] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 32 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0080] The processor 32 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0081] The memory 31 can be used to store the computer programs and / or modules. The processor 32 implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory 31 and calling the data stored in the memory 31. The memory 31 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 31 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0082] It should be noted that the aforementioned electronic devices include, but are not limited to, processors and memory, as will be understood by those skilled in the art. Figure 3 The structural diagram is merely an example of the electronic device described above and does not constitute a limitation on the electronic device. It may include more components than shown in the diagram, or combine certain components, or use different components.
[0083] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed, implements the automated compliance verification method for data elements described in any of the above embodiments.
[0084] It should be understood that the implementation of all or part of the processes in the above-described automated compliance verification method for data elements can also be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described automated compliance verification method for data elements. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0085] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. It should be noted that, for those skilled in the art, several equivalent obvious modifications and / or equivalent substitutions can be made without departing from the technical principles of the present invention, and these obvious modifications and / or equivalent substitutions should also be considered within the scope of protection of the present invention.
Claims
1. A method for automated compliance verification of data elements, characterized in that, include: An industry compliance graph is constructed based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; The data elements to be verified are modeled throughout their entire lifecycle, and metadata is extracted from each processing step. Match the target compliance rules applicable to the metadata from the industry compliance map; Based on the target compliance rules, automated compliance verification is performed on the data elements, and the verification results are output.
2. The automated compliance verification method for data elements as described in claim 1, characterized in that, The construction of the industry compliance map based on multi-source compliance knowledge includes: Compliance text data is collected from multiple channels and preprocessed. A large language model is used to extract rule elements from the preprocessed compliant text data; Identify rule entities from the extracted rule elements and determine the relationships between the rule entities; By using the rule entities as nodes and the relationships as edges, an industry compliance graph is constructed.
3. The automated compliance verification method for data elements as described in claim 1, characterized in that, The data elements to be verified are modeled throughout their entire lifecycle, and metadata for each processing stage is extracted, including: The data element lifecycle is divided into five stages: data collection, data storage, data processing, data circulation, and data application. Metadata management tools are used to extract metadata for each stage of the lifecycle; Based on the extracted metadata, a data lineage relationship is constructed to obtain the complete lineage link of data from the source to the application.
4. The automated compliance verification method for data elements as described in claim 1, characterized in that, The process of matching target compliance rules applicable to the metadata from the industry compliance graph includes: For the metadata of each step, based on the data type, processing scenario and responsible entity of the metadata, calculate the semantic similarity with each compliance rule; The compliance rules whose semantic similarity exceeds a preset threshold are used as the initial compliance rules; Based on the node relationships in the industry compliance graph, a graph neural network is used to perform deep reasoning on the initial compliance rules to identify rule conflicts, rule redundancy, and hidden risks, thereby obtaining the final target compliance rules.
5. The automated compliance verification method for data elements as described in claim 1, characterized in that, The method further includes: A streaming computing framework is used to collect data flow events in real time; these flow events include, but are not limited to, data access events, data modification events, data transmission events, and data sharing events. Perform real-time compliance checks on each process flow event and calculate a risk score; Based on the risk score and the preset multi-level warning thresholds, the corresponding warning mechanism is triggered.
6. The automated compliance verification method for data elements as described in claim 1, characterized in that, The method further includes: Construct a compliance intelligent question-answering system based on RAG retrieval enhancement, and use the industry compliance graph as the knowledge base of the compliance intelligent question-answering system; When a compliance inquiry is received from a user, the system retrieves information from the knowledge base, inputs the retrieved knowledge into the large language model, and outputs the corresponding answer to the compliance question.
7. The automated compliance verification method for data elements as described in claim 6, characterized in that, The method further includes: A risk prediction model is constructed and trained based on time series forecasting algorithms and machine learning classification algorithms. The current compliance risk score, the frequency of historical violations, the trend of data sensitivity changes, the updates of laws and regulations, and the changes in business scenarios are used as input features and input into the risk prediction model. The model outputs the risk probability and risk level predicted by the model for a future preset period. Proactive risk warnings are issued based on the risk probability and the risk level.
8. An automated compliance verification device for data elements, characterized in that, include: The compliance graph construction module is used to construct an industry compliance graph based on multi-source compliance knowledge; wherein, the industry compliance graph uses rule entities as nodes and the relationships between rule entities as edges; The metadata management module is used to perform full lifecycle process modeling of the data elements to be verified and extract metadata for each processing stage. The compliance rule matching module is used to match target compliance rules applicable to the metadata from the industry compliance map. The compliance verification module is used to perform automated compliance verification on data elements based on the target compliance rules and output the verification results.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program; Wherein, when the processor executes the computer program, it implements the data element automated compliance verification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the automated compliance verification method for data elements as described in any one of claims 1 to 7.