Intelligent data management method and system and storage medium

Through technical means such as multimodal artificial intelligence models and generative adversarial networks, the problems of low automation and insufficient privacy protection in traditional data governance have been solved, and the identification and real-time dynamic classification of sensitive information in unstructured data have been achieved, thereby improving the adaptability and real-time nature of data governance.

CN120804080APending Publication Date: 2025-10-17GUANGZHOU SAIBAO LIANRUI INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510949149.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing data governance methods have low levels of automation, lack of real-time and adaptability, are difficult to process unstructured data, have poor privacy protection effects, and cannot meet dynamic regulatory requirements.

Method used

A multimodal artificial intelligence model is used to identify sensitive information, generative adversarial networks and reinforcement learning are combined for data repair, a real-time tracking module is deployed for privacy compliance monitoring, and governance strategies are dynamically updated through a federated learning framework.

Benefits of technology

It realizes the sensitive information identification and real-time dynamic classification of unstructured data, improves data repair efficiency and privacy protection capabilities, dynamically adjusts governance strategies to adapt to changes in the data environment, and improves the adaptability and real-time nature of data governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804080A_ABST
    Figure CN120804080A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, and discloses an intelligent data management method and system and a storage medium, and the method comprises a dynamic data classification and grading step, an intelligent quality restoration step, a privacy compliance real-time monitoring step and a strategy self-optimization step. The system and the storage medium both correspond to the method. By adoption of the data management method and apparatus, the problems of dependence on manual rules, low automation degree, insufficient privacy protection and strategy staticization in traditional data management are solved, full-life-cycle intelligent management of data management is realized, and the data security, compliance and management efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of data governance, and particularly relates to a data intelligent governance method, a system and a storage medium. BACKGROUND

[0002] Data governance refers to the standardized management of the full life cycle (collection, storage, use and destruction) of data through systems, technologies and tools to ensure the accuracy, security, compliance and usability of data. Its core goals include: repairing missing, incorrect or inconsistent data to improve data credibility, meeting the constraints of laws and regulations (such as GDPR and CCPA) on data processing, and dynamically adjusting data governance rules to adapt to changes in the data environment.

[0003] However, the existing data governance method relying on a rule engine defined by artificial (such as sensitive information identification based on regular expressions) has the problems of low efficiency and difficulty in adapting to a dynamically changing data environment, and is only applicable to structured data scenarios and cannot process unstructured data (such as text and images), which leads to difficulty in dealing with complex scenarios (such as sensitive information identification in multi-language mixed text) and limited governance effect. In addition, privacy protection relies on manual configuration and cannot be dynamically triggered according to real-time scenarios, resulting in poor privacy protection effect. The data governance method with the aid of a single machine learning model only applies machine learning technology (such as anomaly detection) to a local link (such as quality repair) of data governance, lacks an intelligent closed loop of the whole process, and lacks real-time monitoring capability, cannot actively intervene in high-risk operations, and leads to high risk of privacy leakage, delayed compliance inspection and difficulty in meeting dynamic supervision requirements.

[0004] Therefore, there is an urgent need for a new data governance technology to solve the deficiencies of traditional data governance in automation, real-time performance and adaptability. SUMMARY

[0005] The application aims to provide a data intelligent governance method, a system and a storage medium to solve the technical problems in the background art and realize intelligent management of the full life cycle of data governance, improve data security, compliance and governance efficiency.

[0006] To achieve the above-mentioned purpose, the application discloses the following technical solutions:

[0007] In a first aspect, the application discloses a data intelligent governance method, which comprises the following steps:

[0008] Dynamic data classification and grading: sensitive information is identified in unstructured data by a multi-modal artificial intelligence model, and data security levels are dynamically divided based on the identified sensitive information through semantic association analysis;

[0009] Intelligent quality repair: generative filling of missing or incorrect data based on a generative adversarial network, and optimization of the repair strategy through reinforcement learning;

[0010] Real-time privacy compliance monitoring: deploying a real-time tracking module in the data flow path, and automatically triggering at least one of desensitization, encryption or access interception operations based on the identified sensitive information through a compliance knowledge base;

[0011] Policy self-optimization: aggregating multi-node governance logs through a federated learning framework, dynamically updating global governance strategies based on the distribution changes of sensitive information and governance effects, and distributing them to each local node for execution.

[0012] As a preferred, the sensitive information identification of unstructured data through a multi-modal artificial intelligence model comprises:

[0013] Identifying sensitive information in unstructured data through a natural language processing model and an image recognition model, the unstructured data including text and images.

[0014] As a preferred, the natural language processing model includes a pre-trained language model and a conditional random field joint model for extracting sensitive information in text; the image recognition model includes a convolutional neural network for identifying sensitive information in images.

[0015] As a preferred, the dynamic division of data security levels based on the identified sensitive information through semantic association analysis comprises:

[0016] Based on the constructed multi-relation knowledge graph, embedding representation of data nodes is performed through a graph neural network, and the association strength with the preset security level is calculated, wherein the multi-relation knowledge graph includes data attributes, context semantics and compliance rules;

[0017] The data security level is dynamically divided according to the association strength.

[0018] As a preferred, the loss function of the generative adversarial network comprises:

[0019] Generator loss, configured to: based on the Wasserstein distance constraint, enhance the diversity and authenticity of the generated data;

[0020] Discriminator loss, configured to: based on the attention mechanism, focus on the semantic consistency of the repair area and the context.

[0021] As a preferred, the aggregation of multi-node governance logs through a federated learning framework comprises:

[0022] Through the federated learning framework, the compliance operation logs of multiple nodes are aggregated, and the semantic rule weights in the compliance knowledge base are dynamically updated, and the adjustment of the weights is based on the risk score distribution;

[0023] The uploaded compliance operation log is added with noise through differential privacy technology to protect data privacy in cross-domain cooperation.

[0024] As preferred, the real-time tracking module deployed in the data flow path automatically triggers the desensitization, encryption or access interception operation based on the identified sensitive information through the compliance knowledge base, including:

[0025] The clauses of the privacy protection regulations are converted into semantic rules in the form of RDF triples to form a compliance knowledge base.

[0026] An edge computing node is deployed in the data flow path to track the identified sensitive information in real time.

[0027] The data operation behavior and the semantic rules in the compliance knowledge base are matched through a graph neural network to generate a risk score, which is a comprehensive evaluation based on the sensitive information type, operation scenario and data magnitude.

[0028] If the risk score exceeds a preset threshold, at least one of the desensitization, encryption or access interception operation is triggered.

[0029] As preferred, the global governance strategy is dynamically updated based on the distribution change of the sensitive information and the governance effect, and is distributed to each local node for execution, including:

[0030] Governance logs are collected from each local node to extract data distribution change features and strategy execution effect indicators.

[0031] A Q-learning model is constructed, taking the current governance strategy as the state and the strategy adjustment action as the input, calculating the reward function based on the extracted features, and generating a strategy update suggestion through iterative training.

[0032] The updated strategy version and the change basis are packaged, the hash value of the strategy version is recorded through a blockchain consensus mechanism, and is synchronized to all local nodes. The local nodes automatically download and execute the updated governance strategy according to the strategy version recorded by the blockchain.

[0033] In a second aspect, the application discloses a data intelligent governance system, which applies the data intelligent governance method as described above, and the system includes:

[0034] The system includes:

[0035] The data analysis module is configured to identify sensitive information in unstructured data through a multi-modal artificial intelligence model, and dynamically divide data security levels based on the identified sensitive information through semantic association analysis.

[0036] a data repair module configured to perform generative filling on missing or incorrect data based on a generative adversarial network and optimize a repair strategy through reinforcement learning;

[0037] a compliance monitoring module configured to deploy a real-time tracking module in a data flow path, automatically trigger at least one of desensitization, encryption or access interception operation based on identified sensitive information through a compliance knowledge base;

[0038] a strategy optimization module configured to aggregate multi-node governance logs through a federated learning framework, dynamically update a global governance strategy based on distribution changes of sensitive information and governance effects, and distribute the global governance strategy to each local node for execution.

[0039] In a third aspect, the present application discloses a computer-readable storage medium comprising at least one memory and at least one processor; the memory is communicatively connected with the processor, and the memory stores a computer program capable of being executed by the processor; when the computer program is executed, the data intelligent governance method as described above is implemented.

[0040] The data intelligent governance method, system and storage medium of the present application solve the problems of relying on manual rules, low automation, insufficient privacy protection and static strategy in traditional data governance. Compared with the prior art, the present application has the following beneficial effects:

[0041] (1) The multi-modal artificial intelligence model is used to identify sensitive information in unstructured data, and the data security level is dynamically divided based on semantic association analysis, which improves the identification ability of sensitive information in unstructured data, realizes real-time dynamic division of data security level, enhances classification accuracy, and adapts to dynamic changes in data environment;

[0042] (2) The generative adversarial network is used to perform generative filling on missing or incorrect data, and the reinforcement learning is used to optimize the repair strategy, which improves the efficiency and semantic consistency of data repair, reduces the consistency error of repaired data, and makes the generative filling more consistent with the context semantic requirements;

[0043] (3) The real-time tracking module is deployed in the data flow path, and the desensitization, encryption or access interception operation is automatically triggered based on the identified sensitive information through the compliance knowledge base, which realizes real-time monitoring and active intervention of sensitive information, reduces the risk of privacy leakage, and reduces the demand for manual intervention;

[0044] (4) The multi-node governance logs are aggregated through the federated learning framework, the global governance strategy is dynamically updated based on the distribution changes of sensitive information and the governance effects, and the global governance strategy is distributed to each local node for execution, which dynamically adjusts the governance strategy to adapt to the data distribution changes, shortens the strategy update cycle, and improves the adaptability and real-time performance of the governance strategy. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Figure 1 The flowchart of the data intelligent governance method provided by the embodiments of the present application. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0048] In this document, the term "comprising" is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the elements defined by the statement "comprising" do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0049] The embodiments provide a data intelligent governance method as shown in Figure 1 The data intelligent governance method includes a dynamic data classification and grading step, an intelligent quality repair step, a real-time privacy compliance monitoring step and a policy self-optimization step, to solve the problems of relying on manual rules, low automation, insufficient privacy protection and static policy in traditional data governance.

[0050] In detail

[0051] Firstly, the dynamic data classification and grading step specifically includes: identifying sensitive information of unstructured data through a multi-modal artificial intelligence model, and dynamically dividing data security levels based on the identified sensitive information through semantic association analysis.

[0052] In the dynamic data classification and grading step, the sensitive information identification of unstructured data through a multi-modal artificial intelligence model is implemented through the following technical means:

[0053] Sensitive information in unstructured data is identified by a natural language processing model and an image recognition model, wherein the sensitive information refers to data information related to user privacy, including at least one or more of an ID number, a face image, a fingerprint image, a bank card number, or a license plate image.

[0054] Further, the natural language processing model includes a pre-trained language model and a conditional random field joint model for extracting sensitive information in text; and the image recognition model includes a convolutional neural network for identifying sensitive information in images.

[0055] In the dynamic data classification and grading step, the data security level is dynamically divided based on the identified sensitive information through semantic association analysis by the following technical means:

[0056] Based on the constructed multi-relation knowledge graph, the data nodes are embedded and represented by a graph neural network, and the association strength with the preset security level is calculated, wherein the multi-relation knowledge graph includes data attributes, context semantics, and compliance rules.

[0057] The data security level is dynamically divided according to the association strength.

[0058] In the above technical means, the data attributes can be information related to data types (such as text, images), sensitive information categories (such as ID numbers, faces), data sources (such as databases, IoT devices), etc.; the context semantics refers to the context features of the data extracted by natural language processing (such as "ID number appears in the reimbursement form") or the semantic association in the image extracted by image recognition (such as "face appears in the monitoring video"); and the compliance rules refer to the conversion of privacy regulations (such as GDPR, CCPA) clauses into RDF triples (RDF is the English abbreviation of Resource Description Framework, which uses the structure of resource-attribute-value (i.e., triples) to provide a framework container, and defines a set of formal methods through XML, which is the structural basis for machine semantic understanding).

[0059] In addition, the multi-relation knowledge graph is designed as follows:

[0060] Node embedding: data entities (such as "a certain image" and "a certain text") are taken as graph nodes, and low-dimensional vector representation is generated by GraphSAGE (Graph Sample and Aggregate, which is a graph neural network algorithm);

[0061] Relation modeling: For multi-relation edges in the multi-relation knowledge graph (such as "data attribute-compliance rule" and "context-sensitive information"), different types of edge weights are processed using a relation graph convolution network.

[0062] Correlation strength calculation: The data node embedding vector output by GraphSAGE is used to calculate the similarity with the preset security level vector (such as "high risk", "medium risk", and "low risk"). The cosine similarity calculation formula is as follows: where h v is the data node embedding vector, and e l is the preset security level vector.

[0063] Secondly, the basis for dividing the data security level is to set a strength threshold, such as a first threshold of 0.9 and a second threshold of 0.7. When the correlation strength is greater than 0.9, the data corresponding to the data node is divided into high risk; when 0.7≤correlation strength<0.9, the data corresponding to the data node is divided into medium risk; and when the correlation strength is less than 0.7, the data corresponding to the data node is divided into low risk.

[0064] Based on this, the data intelligent governance method of the present embodiment realizes dynamic division of data security levels by combining multi-relation knowledge graphs and graph neural networks, uses semantic association analysis to integrate data attributes, context, and compliance rules, improves classification accuracy, and adapts to changes in the data environment through federated learning and graph neural network iteration.

[0065] Secondly, the intelligent quality repair step is specifically: based on the generative filling of the generative adversarial network for missing or incorrect data, and the repair strategy is optimized through reinforcement learning.

[0066] In the intelligent quality repair step, the loss function of the generative adversarial network includes:

[0067] Generator loss, configured to: based on the Wasserstein distance constraint, enhance the diversity and authenticity of the generated data;

[0068] Discriminator loss, configured to: based on the attention mechanism, focus on the semantic consistency of the repair area and the context.

[0069] The Wasserstein distance, also known as the bulldozer distance, is a metric used to measure the difference between two probability distributions. By calculating the Wasserstein distance between the distribution of real and generated data, it addresses the vanishing gradient and mode collapse problems in traditional generative adversarial networks. In the generator network, the input is missing or erroneous data to be repaired (e.g., text fields or image regions), and the output is the repaired data (e.g., completed text fields or inpainted image regions). Multi-scale feature extraction utilizes a U-Net architecture, extracting both global and local features through an encoder-decoder structure. Random noise is injected into the intermediate layers of the generator to enhance the diversity of the generated data. Context-awareness utilizes a convolutional attention mechanism to capture the dependencies between the repaired region and the context. The discriminator loss dynamically focuses on the relationship between the repaired region and the context through an attention mechanism, improving the ability to assess semantic consistency. In the discriminator network, the input is a mixture of real and generated data, and the output is a binary classification probability (real / generated). The multi-scale discriminator uses multiple levels of discriminators to evaluate the global distribution and local details, respectively. The attention mechanism uses a self-attention mechanism to calculate the global correlation between the repaired region and the context.

[0070] Based on this, the data intelligent governance method of this embodiment improves the diversity and semantic consistency during data repair through the combination of Wasserstein distance constraints and attention mechanism, improves the quality of generated data, and has strong adaptability to scenarios such as privacy protection and data quality management.

[0071] Third, the real-time monitoring step for privacy compliance is as follows: deploying a real-time tracking module in the data flow path, and automatically triggering at least one of desensitization, encryption, or access interception operations through the compliance knowledge base based on the identified sensitive information.

[0072] In the real-time privacy compliance monitoring step, a real-time tracking module is deployed in the data flow path. Based on the identified sensitive information, desensitization, encryption, or access blocking operations are automatically triggered through the compliance knowledge base. This is achieved through the following technical means:

[0073] Using natural language processing technology, we identify entities and extract relationships from the terms of privacy protection regulations, converting them into semantic rules in the form of RDF (resource-attribute-attribute-value) triples. Using the Resource Description Framework, we define semantic rules using XML to form a structured database. We then associate the parsed triples with a multi-relational knowledge graph to create a compliance knowledge base.

[0074] Deploying edge computing nodes in the data flow path to track sensitive information in real time; wherein the edge node architecture includes a hardware layer (deploying lightweight edge devices) and a software layer (running lightweight models for sensitive information identification, and only uploading desensitized data to the cloud); the real-time tracking logic includes: monitoring the data flow path (such as API calls, file transfers) through the data flow processing framework, and labeling the identified sensitive information (such as JSON format metadata). The data flow processing involved includes inputting unstructured data (text, image) or structured data (database table), and outputting data flow containing sensitive information labels;

[0075] Through the graph neural network, the data operation behavior (such as reading, writing, and transferring) is matched with the semantic rules in the compliance knowledge base to generate a risk score, which is a comprehensive evaluation based on sensitive information types (such as face, ID number), operation scenarios (such as cross-border transmission), and data magnitude (such as batch export);

[0076] If the risk score exceeds a preset threshold (such as 90 points, which can be dynamically adjusted through governance strategies), at least one of desensitization, encryption, or access interception operations is triggered. Among them, the desensitization operation specifically refers to the fuzzification processing (such as replacing with "***" or hash value) of sensitive fields (such as name, address); the encryption operation specifically refers to the AES-256 algorithm for encrypting storage or transmission of high-risk data; the access interception operation specifically refers to preventing unauthorized users from accessing sensitive data, and recording operation logs to the blockchain audit module.

[0077] In the above technical means, a feasible risk score calculation formula is:

[0078]

[0079] Among them, BaseScore is the default benchmark score (such as 70 points), which is the initial value of the risk score, used to reflect the benchmark risk level of general operations, which can be a benchmark score determined by experience, for example, less than 70 points is considered low risk, and more than 90 points triggers high-risk operation; C i is the weight coefficient set based on the specific scenario and importance of each rule in the compliance knowledge base. When determining C i , the rule attributes corresponding to the regulations are mapped to the predefined coefficient table, such as the requirement of encrypting cross-border data transmission in GDPR Article 44, which corresponds to C i is set to 1.2 (indicating a high-risk scenario); ω i is the preset weight for different sensitive information types, operation scenarios, or data magnitudes, used to quantify their contribution proportion to the risk score. When determining ω iAt the time, the sensitive information type corresponding to the regulatory clause and its compliance requirements are mapped to a preset weight table, such as GDPR Article 9 explicitly requires higher protection measures for biometric data (such as fingerprints, faces), so its ω i is set to 1.5 (indicating a high-risk type); n is the number of matched compliance rules.

[0080] Based on this, the data intelligent governance method of the embodiment realizes intelligent compliance control of the data flow path through RDF triple semantic rule conversion, edge computing real-time tracking, graph neural network dynamic risk scoring, and blockchain audit linkage.

[0081] Fourthly, the strategy self-optimization step, specifically: aggregating multi-node governance logs through a federated learning framework, dynamically updating global governance strategies based on the distribution changes of sensitive information and governance effects, and distributing them to each local node for execution.

[0082] In the strategy self-optimization step, the multi-node governance logs are aggregated through the following technical means:

[0083] The compliance operation logs of multiple nodes (such as the number of interceptions, desensitization effects) are aggregated through a federated learning framework to dynamically update the semantic rule weights in the compliance knowledge base. The adjustment of the weights is based on the risk score distribution (such as the rule priority of high-frequency high-risk scenarios with a risk score ≥ 90 is increased by 20% to increase attention, or such as the rule weight corresponding to operations with a risk score ≤ 30 is reduced by 10% to avoid redundant rules interference);

[0084] The uploaded compliance operation logs are added with noise through differential privacy technology to ensure data privacy in cross-domain collaboration.

[0085] Further, the central server removes noise and restores the true trend using smoothing techniques (such as moving average) after aggregating multi-node logs.

[0086] Based on this, the data intelligent governance method of the embodiment ensures that the uploaded logs cannot be used to infer the original data through differential privacy technology, adjusts the semantic rule weights dynamically to make the governance strategy fit the real-time risk changes, and supports multi-node collaborative optimization through a federated learning framework without sharing sensitive data, reducing the risk of cross-domain collaboration.

[0087] Secondly, in the strategy self-optimization step, the global governance strategy is dynamically updated based on the distribution changes of sensitive information and the governance effects, and distributed to each local node for execution, through the following technical means:

[0088] Collect governance logs from each local node, extract data distribution change features and strategy execution effect indicators;

[0089] A Q-learning model is constructed, taking the current governance strategy as the state, taking the policy adjustment action as the input, calculating the reward function based on the extracted features, and generating policy update recommendations through iterative training;

[0090] After updating the policy version and the change basis (including the reward function parameters of the Q-learning model and the feature weights), the hash value of the policy version is recorded through the blockchain consensus mechanism (such as practical Byzantine fault tolerance or proof of work), and is synchronized to all local nodes. The local nodes automatically download and execute the updated governance strategy according to the policy version recorded in the blockchain.

[0091] Among them, the governance log includes sensitive information identification results, repair operation records, and compliance operation logs. The sensitive information identification results are the detection results of sensitive fields (such as ID numbers and face images) recorded by the multi-modal model, including field types, confidence levels, and context semantic labels. The repair operation record is the context semantic consistency score of the missing data segment filled by the generative adversarial network and the repair area. The compliance operation log records the number of desensitization operations (such as the number of blurred fields), encryption trigger events (such as the number of AES-256 encryptions), and access interception records (such as intercepted user IDs and interception rule names). The data distribution change feature includes sensitive information type frequency change and data security level distribution offset. The sensitive information type frequency change is the time series distribution of sensitive information types (such as the proportion of "face images" and the frequency of ID numbers) in each local node, and the offset (such as Z-score) from the global baseline is calculated. The data security level distribution offset is the embedding representation of the graph neural network based on the knowledge graph, which analyzes the distribution change of the data node security level (such as "high risk" and "low risk") and identifies abnormal offset scenarios (such as a node suddenly appearing a large amount of high-risk data). The policy execution effect indicator includes desensitization coverage, encryption response delay, and access interception accuracy. The desensitization coverage is the proportion of sensitive fields that are successfully desensitized in the local node (such as "name field desensitization rate"). The encryption response delay is the time interval from detecting high-risk data to completing the encryption operation (such as an average encryption delay of <100ms). The access interception accuracy is verified by the blockchain audit module to ensure the compliance of the interception operation (such as whether the interception rule complies with Article 30 of GDPR).

[0092] The model architecture and input / output of the Q-learning model are as follows:

[0093] State space: current global governance strategy (such as desensitization intensity threshold and encryption algorithm type), extracted feature vector (sensitive information type frequency change, data security level distribution offset, desensitization coverage, etc.);

[0094] Action space: policy adjustment actions are a finite set (e.g., "increase desensitization strength in high-risk scenarios", "increase encryption priority for cross-border transmissions");

[0095] Reward function: compliance score (quantified based on GDPR clause satisfaction), data availability (consistency error of data after repair (e.g., similarity of GAN generated fields to original data <5%)), resource consumption (edge computing node load (e.g., CPU utilization, memory occupancy)), and the reward function is derived by weighting the compliance score, data availability, and resource consumption.

[0096] The training of the Q-learning model and the generation of the policy are as follows:

[0097] Initialize the parameters of the Q-learning model;

[0098] Input the current state (governance policy + feature vector), and select a policy adjustment action (e.g., "increase desensitization strength");

[0099] Calculate the reward value;

[0100] Repeat the iteration until convergence;

[0101] Output the optimal action sequence (e.g., "increase desensitization strength in high-risk scenarios", "decrease encryption priority for low-risk data") as a policy update recommendation.

[0102] Based on this, the data intelligent governance method of the embodiment realizes the minute-level dynamic update and safe distribution of the global governance policy through the combination of federated learning log aggregation, Q-learning policy optimization, and blockchain synchronization mechanism.

[0103] The embodiment provides a data intelligent governance system in a second aspect, which applies the data intelligent governance method as described above, and the system comprises:

[0104] A data analysis module configured to: identify sensitive information from unstructured data through a multi-modal artificial intelligence model, and dynamically divide data security levels based on the identified sensitive information through semantic association analysis;

[0105] A data repair module configured to: perform generative filling of missing or incorrect data based on a generative adversarial network, and optimize the repair strategy through reinforcement learning;

[0106] A compliance monitoring module configured to: deploy a real-time tracking module in the data flow path, and automatically trigger at least one of desensitization, encryption, or access interception operations based on the identified sensitive information through a compliance knowledge base;

[0107] The policy optimization module is configured to aggregate the multi-node governance logs through a federated learning framework, dynamically update a global governance policy based on distribution changes of sensitive information and governance effects, and distribute the global governance policy to each local node for execution.

[0108] It should be noted that the data intelligent governance system of the present embodiment corresponds to the foregoing data intelligent governance system, and therefore, the parts not disclosed in detail in the present text (including but not limited to technical effects and specific implementation technical means) can be correspondingly referred to the related description in the foregoing data intelligent governance method, which will not be repeated herein.

[0109] The present embodiment provides a computer readable storage medium in a third aspect, comprising at least one memory and at least one processor; the memory is communicatively connected with the processor, and the memory stores a computer program capable of being executed by the processor; when the computer program is executed, the data intelligent governance method as described above is realized.

[0110] Similarly, it should be noted that the computer readable storage medium of the present embodiment corresponds to the foregoing data intelligent governance method, and therefore, the parts not disclosed in detail in the present text (including but not limited to technical effects and specific implementation technical means) can be correspondingly referred to the related description in the foregoing data intelligent governance method, which will not be repeated herein.

[0111] In the embodiments provided by the present application, it should be understood that the embodiments described herein can be realized by hardware, software, firmware, middleware, codes or any proper combination thereof. For hardware implementation, the processor can be realized in one or more of the following components: an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, other electronic units designed to perform the functions described herein, or a combination thereof. For software implementation, the procedures described herein can be implemented with a computer program that is written in any suitable programming language. The program can be stored in a computer readable storage medium or transmitted as one or more instructions or codes on the computer readable storage medium. The computer readable storage medium includes any storage medium that can be accessed by a computer. The computer readable storage medium can include but is not limited to the following media: a RAM, a ROM, an EEPROM, a CD-ROM or other optical disc storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer.

[0112] Finally, it should be noted that the above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, modifications or equivalent replacements of some technical features described in the foregoing embodiments can be made by those skilled in the art, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data intelligent governance method, characterized in that: The method comprises the following steps: Dynamic data classification and grading: Use multimodal AI models to identify sensitive information in unstructured data, and dynamically classify data security levels based on the identified sensitive information through semantic association analysis; Intelligent quality repair: Generates missing or erroneous data using a generative adversarial network and optimizes the repair strategy through reinforcement learning; Real-time privacy compliance monitoring: Deploy a real-time tracking module in the data flow path, and automatically trigger at least one of the following actions: desensitization, encryption, or access blocking based on the identified sensitive information through the compliance knowledge base; Strategy self-optimization: Aggregate multi-node governance logs through the federated learning framework, dynamically update the global governance strategy based on the distribution changes of sensitive information and governance effects, and distribute it to each local node for execution.

2. The data intelligent governance method according to claim 1, characterized in that: The identification of sensitive information in unstructured data using a multimodal artificial intelligence model includes: Sensitive information in unstructured data is identified through natural language processing models and image recognition models, where the unstructured data includes text and images.

3. The data intelligent management method according to claim 2, characterized in that: The natural language processing model includes a pre-trained language model and a conditional random field joint model, which is used to extract sensitive information from text; the image recognition model includes a convolutional neural network, which is used to identify sensitive information in images.

4. The data intelligent governance method according to claim 2, characterized in that: The dynamic classification of data security levels based on the identified sensitive information through semantic association analysis includes: Based on the constructed multi-relational knowledge graph, data nodes are embedded and represented through a graph neural network, and their association strength with the preset security level is calculated. The multi-relational knowledge graph includes data attributes, contextual semantics, and compliance rules. Dynamically divide data security levels according to association strength.

5. The data intelligent management method according to claim 1, characterized in that: The loss function of the generative adversarial network includes: Generator loss, configured as follows: Based on the Wasserstein distance constraint, it enhances the diversity and authenticity of the generated data; The discriminator loss is configured as follows: based on the attention mechanism, it focuses on the semantic consistency between the repaired area and the context.

6. The data intelligent governance method according to claim 1, characterized in that: Aggregating multi-node governance logs through the federated learning framework includes: Aggregate compliance operation logs from multiple nodes through a federated learning framework and dynamically update semantic rule weights in the compliance knowledge base. The weights are adjusted based on the risk score distribution. Differential privacy technology is used to add noise to uploaded compliance operation logs to ensure data privacy in cross-domain collaboration.

7. The data intelligent governance method according to claim 1 or 6, characterized in that: The real-time tracking module is deployed in the data flow path to automatically trigger desensitization, encryption, or access blocking operations based on the identified sensitive information through the compliance knowledge base, including: Convert the terms of privacy protection regulations into semantic rules in the form of RDF triples to form a compliance knowledge base; Deploy edge computing nodes in the data flow path to track the identified sensitive information in real time; Using graph neural networks, data manipulation behaviors are matched with semantic rules in the compliance knowledge base to generate a risk score, which is a comprehensive assessment based on the type of sensitive information, the operation scenario, and the amount of data. If the risk score exceeds the preset threshold, at least one of desensitization, encryption or access interception operations is triggered.

8. The data intelligent governance method according to claim 1 or 6, characterized in that: The global governance strategy is dynamically updated based on the distribution changes and governance effects of sensitive information and distributed to each local node for execution, including: Collect governance logs from each local node and extract data distribution change characteristics and policy execution effect indicators; Build a Q-learning model, using the current governance policy as the state and the policy adjustment action as the input. It calculates the reward function based on the extracted features and generates policy update recommendations through iterative training. After packaging the updated policy version and the basis for the changes, the hash value of the policy version is recorded through the blockchain consensus mechanism and synchronized to all local nodes. The local nodes automatically download and execute the updated governance policy based on the policy version recorded on the blockchain.

9. A data intelligent management system, applying the data intelligent management method according to any one of claims 1 to 8, characterized in that: The system includes: The data analysis module is configured to: identify sensitive information in unstructured data through a multimodal artificial intelligence model, and dynamically classify data security levels based on the identified sensitive information through semantic association analysis; The data repair module is configured to: generatively fill in missing or erroneous data based on a generative adversarial network and optimize the repair strategy through reinforcement learning; A compliance monitoring module is configured to: deploy a real-time tracking module in a data flow path, and automatically trigger at least one of desensitization, encryption, or access blocking operations based on the identified sensitive information through a compliance knowledge base; The policy optimization module is configured to aggregate multi-node governance logs through the federated learning framework, dynamically update the global governance policy based on the distribution changes of sensitive information and the governance effect, and distribute it to each local node for execution.

10. A computer-readable storage medium, characterized in that It includes at least one memory and at least one processor; the memory is communicatively connected to the processor, and the memory stores a computer program that can be executed by the processor. When the computer program is executed, the data intelligent governance method as described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Data management method and device, computer equipment and storage medium

    CN121743312A