Data management system and method based on AI large model

By using an AI-based big data model-based data governance system, the shortcomings of existing data governance tools in terms of intelligent recommendation, deep insight, and user experience are addressed. This system enables intelligent and efficient management of the entire data governance process, improves the real-time performance and security of data processing, and is applicable to data-driven decision-making across multiple industries.

CN121858652APending Publication Date: 2026-04-14SHAANXI HANGUANG DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI HANGUANG DIGITAL TECH CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing data governance tools rely on human experience for governance rule configuration, lack scenario-based intelligent recommendation capabilities, and lack in-depth insights and automated interpretation of governance results. The system operation is complex and the user experience is unfriendly, resulting in low configuration efficiency, high threshold, and difficulty in achieving intelligence and widespread application.

Method used

The data governance system based on AI big data models includes a data integration and access module, a core data processing and governance module, an AI big data model intelligent engine module, a data asset and service module, a user interaction and intelligent assistance module, and a system management and support module. It collaborates through a unified service bus and API interface, and uses the AI ​​big data model intelligent engine module to provide intelligent recommendation, automated configuration, in-depth interpretation, and interactive guidance capabilities.

Benefits of technology

It achieves intelligent data governance throughout the entire process, significantly improving governance efficiency and accuracy, lowering the barrier to entry for system use, enhancing the real-time performance and security of data processing, providing efficient processing and quality closed-loop management of multi-source heterogeneous data, and supporting data-driven decision-making and data security protection in multiple industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858652A_ABST
    Figure CN121858652A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of data governance, and discloses a data governance system and method based on an AI large model, and the system comprises a data integration and access module, a data processing and governance core module, an AI large model intelligent engine module, a data asset and service module, a user interaction and intelligent auxiliary module, and a system management and support module. All the modules carry out data interaction and function collaboration through a unified service bus and an API interface. On the basis of a mature data governance function system, a leading-edge AI large model technology is deeply fused, and full-process intelligent upgrading from data integration, processing, governance and service to interactive guidance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to, but is not limited to, the field of data governance technology, and particularly relates to a data governance system and method based on a large AI model. Background Technology

[0002] As enterprises deepen their digital transformation, data has become a core strategic asset. Effective data governance is fundamental to ensuring data quality, enhancing data value, and guaranteeing data security and compliance. Traditional data governance tools generally adopt a modular design, providing standardized operating interfaces and process configuration capabilities around data integration, metadata management, data quality management, and master data management. For example, the "data governance platform" known in this field provides full-process functional coverage from multi-source data access and source-linked data management to data cleaning, quality rule configuration, data lineage analysis, and asset catalog management and service approval, realizing the tooling and partial automation of data governance tasks.

[0003] However, with the growth of data scale and the increase in business complexity, traditional tools have gradually revealed their shortcomings in areas such as intelligent decision-making, adaptive learning, and user-friendly experience. Although such tools have relatively complete functional systems, they still rely heavily on human experience and pre-set templates in areas such as adaptive recommendation of governance rules, intelligent generation and interpretation of governance reports, and intelligent guidance of system operations, and have not yet achieved true "intelligent governance." In recent years, artificial intelligence technologies, represented by large-scale language models, have made breakthrough progress, demonstrating strong potential in natural language understanding, knowledge reasoning, content generation, and interactive guidance, providing a feasible technical path for data governance to move from "tool-assisted" to "intelligent-driven."

[0004] Based on an analysis of the current state of technology in the field of data governance, especially mature products like "data asset management toolkits," existing technologies still have the following core shortcomings that urgently need to be addressed in the process of evolving towards intelligence:

[0005] (1) The configuration of governance rules relies on human experience and lacks the ability to make intelligent recommendations based on specific scenarios.

[0006] In existing systems, key governance aspects such as setting data quality rules (requiring either selecting built-in rules or writing Python scripts), mapping data standards, and choosing data anonymization strategies all require users to manually configure them based on their own knowledge. The system cannot automatically recommend or generate optimal governance rules based on the characteristics of the data itself, the business context, and historical governance results, leading to low configuration efficiency, high barriers to entry, and difficulty in ensuring the adaptability and effectiveness of rules in different scenarios.

[0007] (2) The output of governance results is mainly in the form of fixed reports, lacking in-depth insights and automated interpretation.

[0008] Existing systems can generate deliverables such as quality inspection logs and defect reports, but these are mostly structured data lists or statistical charts based on fixed templates. The reports lack business context analysis of the root causes of problems, assessment of potential impacts, and personalized interpretations for different audiences (such as business personnel, technical personnel, and managers). This makes it difficult to directly translate governance results into business actions, and the value of the reports is not fully realized.

[0009] (3) The system is complex to operate and lacks intelligent guidance, resulting in limited user experience and empowerment.

[0010] Existing tools are feature-rich and have highly specialized operating procedures, requiring extensive training for new users to master. The system interface lacks intelligent navigation based on user roles and tasks, real-time prompts for operation steps, and natural language interactive question-and-answer functions. This results in high learning costs and low usage efficiency, preventing the system from achieving efficient and widespread accessibility and hindering the widespread establishment and deep implementation of a data governance culture.

[0011] Therefore, this invention aims to build a new generation of data governance system based on the functional system of existing data governance toolkits, deeply integrating AI big model technology, and possessing the capabilities of autonomous analysis, intelligent recommendation, dynamic interpretation, and interactive guidance. This will achieve an intelligent leap forward in the entire data governance process while inheriting the functional completeness advantages of existing tool platforms.

[0012] Based on the above analysis, the urgent technical problems that need to be solved in the existing technology are:

[0013] (1) How to reduce the reliance on human experience in configuring data governance rules and achieve intelligent recommendation and automated generation of governance strategies;

[0014] (2) How to transform structured quality inspection results into intelligent assessment reports with in-depth business insights and actionable recommendations;

[0015] (3) How to lower the threshold for using the system and provide users with intelligent operation guidance and interactive assistance. Summary of the Invention

[0016] To address the problems existing in the prior art, this invention provides a data governance system and method based on a large AI model.

[0017] This invention is implemented as follows: a data governance system based on an AI large model, characterized in that the system includes a data integration and access module, a data processing and governance core module, an AI large model intelligent engine module, a data asset and service module, a user interaction and intelligent assistance module, and a system management and support module. Each module interacts with data and collaborates with functions through a unified service bus and API interface.

[0018] Furthermore, the data integration and access module, serving as the entry point for enterprise data assets into this system, is responsible for connecting with various customer business systems to complete data extraction, de-identification, and initial data entry into the data lake. It adopts a layered design, with the following specific structure:

[0019] 1) Multi-source data connection management layer: Provides a unified configuration interface, supporting access to various heterogeneous data sources including but not limited to MySQL, Oracle, SQL Server, PostgreSQL, DB2, DM, Kingbase, TiDB, MongoDB, Redis, Hive, HBase, and Kafka. Users can input connection information such as host address, port, database name, username, and password through a graphical interface. This layer provides connection testing functions and registers and manages successfully connected data sources. It allows users to view the list of databases and tables under the data source and execute simple SQL probing. Its core capability is to shield the differences of the underlying data sources and provide standardized data access capabilities for the upper layer.

[0020] 2) Intelligent Data Masking Layer: Provides configurable data masking capabilities. Through the task configuration interface, users first select the registered source data source and the specific data table to be masked. The system then displays all fields of the table, from which users can select fields containing sensitive information such as personal privacy and trade secrets. This layer has a rich and extensible masking rule library built-in, with rule types covering encryption, replacement, simulation, masking, and hashing.

[0021] 3) Data Extraction and Integration Layer: This layer is responsible for synchronizing the source data stably, efficiently, and in full or incremental form to the system's data storage core. This layer provides two integration modes: fast data integration mode and professional data integration mode.

[0022] 4) Source Data Storage and Management Layer: As the entry layer of the system data lake (ODS, Operational DataStore), persistent storage accesses the original or near-original data through the above integration method. This layer not only provides storage, but also provides powerful management and exploration functions, including table structure and metadata query, data preview, and end-to-end access log monitoring.

[0023] Furthermore, the core module for data processing and governance is the core of this system's data standardization and normalization. It performs a series of governance operations on the source data after it enters the lake, transforming it into high-quality, reliable, and usable data assets, including:

[0024] 1) Metadata Management Submodule: Used to break down data silos, build a unified data asset map for the enterprise, support the creation of collection tasks, provide a tree-shaped directory management interface, and allow users to create multi-level directory structures based on business themes, departmental functions, and other dimensions;

[0025] 2) Data Governance Execution Submodule: Used to implement specific governance measures to ensure data consistency and standardization, including data lineage analysis, master data management, and data tag management;

[0026] 3) Data Quality Closed-Loop Management Submodule: Used to measure and improve the intrinsic quality of data, forming a management closed loop, and can perform quality rule management, quality inspection and evaluation, and intelligent quality enhancement.

[0027] Furthermore, the AI ​​large-scale model intelligent engine module is an independently deployed intelligent computing center that provides services externally via API. Its core lies in using a large language model finely tuned with domain knowledge to inject three core AI capabilities into all the aforementioned functional modules:

[0028] 1) Intelligent recommendation and automated configuration generation: For configuration scenarios that require professional knowledge, such as data integration and access modules (selection of de-identification rules) and core data processing and governance modules (quality rules, cleaning rules, data standard mapping configuration), when a user makes a request, the engine receives relevant context such as data schema, samples, and business goal descriptions, and outputs the optimal configuration recommendation through model reasoning;

[0029] 2) Intelligent report generation and in-depth interpretation: The engine receives structured results input, combines rich business metadata, and uses natural language generation and logical reasoning technology to automatically generate structured and narrative analysis reports;

[0030] 3) Intelligent Interaction and Context-Aware Guidance: Serving as the driving engine for user interaction and intelligent assistance modules. It understands the questions posed by users through natural language, retrieves and organizes information from the system knowledge base and user manual, and provides clear, step-by-step solutions.

[0031] Furthermore, the data asset and service module, geared towards data consumers, encapsulates, securely publishes, and conveniently delivers governed, high-quality data assets. This module includes:

[0032] 1) Data Asset Catalog: An asset portal for business users is built based on the metadata management submodule;

[0033] 2) Data as a service: Supports encapsulating data query and processing logic into standardized, reusable API services;

[0034] 3) Service Operations and Security Gateway: Build a secure access and control system for data assets.

[0035] Furthermore, the user interaction and intelligent assistance module is a unified front-end interface for the system to interact with various types of users, including:

[0036] 1) Traditional Functional Interface: Provides a clear and responsive web interface, with navigation via the left-hand menu tree to major functional areas such as data integration, data governance, data quality, data services, and system management;

[0037] 2) Embedded intelligent assistance system: A smart assistant dialogue floating window is permanently embedded in the system interface, allowing users to ask questions in natural language at any time;

[0038] 3) Contextualized intelligent guidance and prompts: When users are performing critical and complex tasks, the system can sense the current step the user is in and provide intelligent guidance dynamically and non-intrusively.

[0039] Furthermore, the system management and support module provides infrastructure, management tools, and operational support for the entire data governance system, including:

[0040] 1) Basic Resource Management: Provides multi-level organizational structure management for institutions, positions, and users, enabling the mapping between users and entity organizations;

[0041] 2) System Configuration and Operation: Maintain the standard code set such as enumeration values ​​and status codes used globally by the system, uniformly manage all jobs that need to be executed periodically, uniformly store and manage user-uploaded documents, images and other attachment resources, centrally display the running status, resource consumption and success / failure records of all batch processing jobs and stream processing jobs, and provide operation and maintenance operation entry points such as job start and stop, log viewing and alarm configuration, so as to realize transparent supervision of system operation status.

[0042] Another objective of this invention is to provide a data governance method based on large AI models, which specifically includes:

[0043] S1: Governance Task Intent Reception and Intelligent Parsing: The system receives a data governance task request and calls the task understanding and decomposition service of the AI ​​large model intelligent engine module;

[0044] S2: Governance Context Awareness and Resource Discovery: The process automatically triggers a context-aware agent, which autonomously calls the APIs of multiple modules within the system to collect and integrate key information;

[0045] S3: Dynamic Task Planning and Intelligent Agent Collaborative Orchestration: The AI ​​large model intelligent engine module dynamically generates an optimal "governance execution plan" based on the set of sub-goals parsed in step S1 and the context information collected in step S2.

[0046] S4: Modular execution and closed-loop feedback driven by intelligent agents: Each professional intelligent agent is instantiated and begins execution according to the planned workflow;

[0047] S5: Governance Results Aggregation and Adaptive Delivery: When all agents complete their tasks or the process reaches the predetermined termination conditions, the results aggregation phase begins. A results aggregation agent collects all key products generated in the entire intelligent governance process. The process delivers the results in the most appropriate way based on the initial source of the task request.

[0048] S6: Knowledge Accumulation and Model Iteration Optimization: The system accumulates the execution records of each complete process, context information, AI decision points, and final user feedback as training data. This data can be periodically used to fine-tune or optimize the core model in the AI ​​large model intelligent engine module through reinforcement learning.

[0049] Furthermore, in S2, the key information includes:

[0050] 1) Data Asset Context: From the metadata management submodule, query all data tables, fields, and their business meanings related to "customer";

[0051] 2) Technical Status Context: Explore the current status, access method, update frequency, and historical data quality overview of relevant data tables from the source data storage and management layer and the data integration and access module;

[0052] 3) Business Relationship Context: Obtain the lineage link between "customer" related data and downstream "marketing analysis reports" from the data lineage analysis function;

[0053] 4) Rule knowledge context: Retrieve existing, reusable quality rules and desensitization rule templates from the quality rule management and intelligent data desensitization layer;

[0054] 5) Environment and Constraint Context: Consider the current system load, available computing resources, and compliance policy requirements.

[0055] Furthermore, in S5, a result aggregation intelligence system collects all key outputs generated during the entire intelligent governance process, including: executed task logs, newly created or modified governance rule configurations, generated data quality scores, intelligent evaluation reports, etc.; finally, the process delivers the results in the most appropriate manner based on the initial source of the task request.

[0056] 1) For requests from the intelligent assistant, the results will be presented to the user in the form of a structured natural language summary in the dialogue interface, along with a link to a detailed report;

[0057] 2) For triggers from function buttons, the system will automatically navigate the user to a comprehensive results dashboard page, which provides a panoramic view of the governance process and results;

[0058] 3) For tasks that are triggered on a schedule, the results will be recorded and may trigger notifications or be included in the daily operation and maintenance report.

[0059] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0060] First, it significantly improves the automation and intelligence of data governance: By introducing an AI large-scale model core engine module, combined with dynamic rule configuration and online learning capabilities, this invention achieves automation and intelligence in data cleaning, quality monitoring, anomaly detection, and repair. Traditional technologies rely on manual intervention and predefined rules, which are inefficient and prone to errors. In contrast, this invention achieves full-process automation through AI technology, greatly improving the efficiency of data governance, reducing the need for manual intervention, and enhancing the accuracy and reliability of data processing.

[0061] Enhancing the real-time and dynamic nature of data processing: By employing a real-time data stream processing module and utilizing streaming computing engines such as Apache Flink or Spark Streaming, combined with a dynamic policy adjuster, this invention enables efficient processing and analysis of real-time data streams and dynamically adjusts governance strategies based on real-time data quality feedback. Traditional technologies typically employ batch processing methods, which cannot meet real-time requirements. This invention, however, achieves rapid response and dynamic optimization of real-time data streams, ensuring the real-time performance and accuracy of data in dynamic environments, thus meeting the high real-time requirements of big data environments.

[0062] Optimizing Multi-Source Heterogeneous Data Processing and Data Quality Closed-Loop Management: This invention achieves comprehensive access and standardized processing of multi-source heterogeneous data through a data integration and preprocessing module. Combined with the data fusion and matching capabilities of the AI ​​large-scale model core engine module, a unified data view is formed. Simultaneously, the data quality closed-loop management module enables full-cycle management from quality issue discovery to remediation. For example, in the data integration and preprocessing module, the system can comprehensively access various types of multi-source heterogeneous data from enterprises within the park, including business registration data, patent and honor data, financial data, and operational data, and perform efficient cleaning and standardization. In the data quality closed-loop management module, the system is responsible for completing the full-cycle management of data quality from issue discovery to quality remediation, ensuring high data quality throughout its entire lifecycle and providing reliable data support for park managers. This addresses the shortcomings of existing technologies in multi-source heterogeneous data processing and data quality closed-loop management.

[0063] Strengthening Data Security and Privacy Protection Capabilities: This invention employs homomorphic encryption and differential privacy technologies to protect sensitive data through a data security and privacy protection module, and identifies and handles security risks in real time through a security risk detection layer. Simultaneously, the prediction and optimization layer of the AI ​​large-scale model core engine module predicts future data quality based on historical data and quality trends, and optimizes governance strategies. For example, in the data security and privacy protection module, the system uses advanced encryption technology to encrypt the storage and transmission of sensitive data, protecting user privacy. In the AI ​​large-scale model core engine module, the system can predict future data quality based on historical data and quality trends, proactively implementing optimized governance strategies to achieve intelligent and forward-looking data governance, improving data security and governance effectiveness, and addressing the shortcomings of existing technologies in data security and AI-driven predictive optimization capabilities.

[0064] Secondly, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:

[0065] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:

[0066] This invention's technical solution deeply integrates AI large-scale models with data governance processes, providing enterprises with end-to-end intelligent support from data access and governance to service provision. Its implementation will significantly reduce the human and technical barriers for enterprises in the field of data governance, improve the availability, credibility, and compliance of data assets, thereby accelerating the implementation of data-driven decision-making. At the business level, this system can form a business model that combines standardized products and customized solutions, providing SaaS services or localized deployments to data-dependent industries such as finance, healthcare, government affairs, and manufacturing, helping enterprises build a self-controllable data intelligence foundation. It is expected to bring the following business value:

[0067] Reduce operating costs: By using AI-powered intelligent recommendations and automated configuration, reduce reliance on professional personnel in data governance, shorten data preparation cycles, and lower long-term operation and maintenance costs.

[0068] Enhancing the efficiency of data value transformation: Intelligent reporting and in-depth interpretation functions enable governance results to directly serve business decisions, improving the application effectiveness of data assets in business innovation, risk management, customer insights, and other aspects.

[0069] Enhance market competitiveness: The system supports real-time governance and quality closed-loop management of multi-source heterogeneous data, which can help enterprises respond quickly to market changes, build a compliant and reliable data supply chain, and gain an advantage in digital transformation.

[0070] Promoting the construction of a governance ecosystem: It can serve as a core component of an enterprise data platform, promoting the orderly integration and secure circulation of internal and external data resources, and providing technical support for the marketization of data elements. Attached Figure Description

[0071] Figure 1 This is a schematic diagram of the overall architecture of the data governance system based on AI large model provided in an embodiment of the present invention;

[0072] Figure 2 This is a flowchart illustrating the data governance algorithm based on a large AI model provided in an embodiment of the present invention;

[0073] Figure 3 This is a schematic diagram of the composition of the data integration and access module provided in an embodiment of the present invention;

[0074] Figure 4 This is a schematic diagram illustrating the structure and interaction of the AI ​​large-scale model intelligent engine module provided in an embodiment of the present invention;

[0075] Figure 5 This is a schematic diagram of the interface layout and guidance process of the user interaction and intelligent assistance module provided in an embodiment of the present invention;

[0076] Figure 6 This is a schematic diagram of the quality evaluation execution process provided in an embodiment of the present invention.

[0077] Figure 7 This is a schematic diagram of the data quality evaluation report content provided in an embodiment of the present invention. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0079] like Figure 1As shown in the figure, this invention provides a data governance system based on an AI large-scale model. This system adopts a modular and layered architecture design. While retaining mature full-process data governance functions, it deeply integrates an AI large-scale model as the intelligent core, achieving intelligent, automated, and user-friendly interaction in the governance process. The system mainly includes the following modules: a data integration and access module, a data processing and governance core module, an AI large-scale model intelligent engine module, a data asset and service module, a user interaction and intelligent assistance module, and a system management and support module. Each module interacts with data and collaborates functionally through a unified service bus and API interface.

[0080] (1) Data integration and access module

[0081] This module serves as the entry point for enterprise data assets into this system. It is responsible for connecting with various customer business systems to complete data extraction, anonymization, and initial data import. It adopts a layered design, such as... Figure 3 As shown, the specific structure is described in detail below:

[0082] 1) Multi-Source Data Connection Management Layer: Provides a unified configuration interface, supporting access to various heterogeneous data sources including but not limited to MySQL, Oracle, SQL Server, PostgreSQL, DB2, DM, Kingbase, TiDB, MongoDB, Redis, Hive, HBase, and Kafka. Users input connection information such as host address, port, database name, username, and password through a graphical interface. This layer provides connection testing functionality and registers and manages successfully connected data sources, allowing users to view the list of databases and tables under the data source and execute simple SQL probing queries. Its core capability lies in shielding the differences in underlying data sources, providing standardized data access capabilities to the upper layers.

[0083] 2) Intelligent Data Masking Layer: To ensure the security and compliance of sensitive data during the integration process, this layer provides configurable data masking capabilities. Users first select the registered source data source and the specific data table to be masked through the task configuration interface. The system then displays all fields of the table, from which users can select fields containing sensitive information such as personal privacy, trade secrets, etc., such as name, ID number, mobile phone number, address, email, bank card number, etc.

[0084] This layer includes a rich and extensible de-identification rule library, covering rule types such as encryption, replacement, simulation, masking, and hashing. Encryption uses reversible encryption algorithms such as AES; replacement uses fixed or ambiguous values, such as "*"; simulation generates non-real data that conforms to real rules, such as virtual ID numbers; masking retains some features, such as displaying only the first three and last four digits of a phone number; hashing uses algorithms such as MD5 and SHA-256 for hash processing.

[0085] Its intelligent features are reflected in the following: when a user selects a field such as "ID number," the intelligent recommendation service of the AI ​​large-scale model intelligent engine module can be triggered. This service automatically analyzes and highlights the most suitable de-identification rules based on field names, data samples, industry compliance requirements such as the "Personal Information Protection Law," and best practices. For example, for ID numbers, it recommends the "mask" rule, which users can adopt with one click or manually adjust. The de-identification task can be configured for "one-time" execution or "periodic" automatic execution as set by the CRON expression. After execution, a de-identified data table with the suffix "_masking" will be generated on the target end, thus achieving secure data storage.

[0086] 3) Data Extraction and Integration Layer: This layer is responsible for stably, efficiently, and fully or incrementally synchronizing source data to the system's data storage core. The source data includes raw data or secure data processed by the anonymization layer. To adapt to different business scenarios and technical requirements, this layer provides two integration modes:

[0087] Rapid Data Integration Mode: Suitable for table-to-table synchronization scenarios with simple structures and moderate timeliness requirements. Users can quickly create a data integration pipeline by completing the following steps through a wizard-driven interface: "Basic Information Configuration" (task name, scheduling type), "Synchronization Source Configuration" (selecting the data source and table or custom SQL), "Target Configuration" (specifying the storage location, typically the ODS layer, and configuring the processing strategy when the table exists), and "Mapping and Transformation Configuration." This mode is ready to use out of the box, greatly reducing the difficulty of basic data synchronization.

[0088] Professional Data Integration Mode: Suitable for handling complex data transformations, high-performance synchronization, real-time streaming access, or fine-grained incremental synchronization scenarios. This mode is deeply encapsulated and enhanced based on open-source high-performance data integration frameworks such as Apache SeaTunnel. Users can write or configure more complex extraction, transformation, and loading logic through a more professional configuration interface. It supports full synchronization, incremental synchronization based on timestamps or incremental fields, and provides dirty data control, task resource allocation, and advanced scheduling strategies. This mode provides data engineers with powerful flexibility to meet the stringent data integration needs of enterprises.

[0089] 4) Source Data Storage and Management Layer: As the entry layer of the system's data lake (ODS, Operational DataStore), persistent storage accesses raw or near-raw data through the integration methods described above. This layer not only provides storage but also offers powerful management and exploration capabilities, including:

[0090] ① Table Structure and Meta-information Query: Displays detailed information of all access tables in a structured manner, such as table name, field name, data type, length, precision, whether it is nullable, default value, primary key, field comments, table comments, etc., to help users quickly understand the data structure.

[0091] ② Data Preview: Allows users to perform quick data sampling on any table, such as the first 1000 rows or random sampling, and to intuitively view the data content, format and general quality without writing SQL, providing a basis for subsequent governance decisions.

[0092] ③ End-to-end access log monitoring: Detailed records of the entire lifecycle of data access for each table, including the data source system, the executed synchronization task ID, synchronization start and end times, the number of synchronized data rows, the data volume, task execution status (success / failure), and error messages (if failure). This ensures complete traceability of the data source, laying a solid foundation for tracing data lineage and pinpointing the source of data quality issues.

[0093] (2) Core module of data processing and governance

[0094] This module is the core of the system for achieving data standardization and normalization. It performs a series of governance operations on the source data after it enters the lake, transforming it into high-quality, reliable, and usable data assets.

[0095] 1) Metadata Management Submodule: Designed to break down data silos and build a unified data asset map for the enterprise.

[0096] ① Metadata Acquisition: Supports the creation of acquisition tasks to automatically scan and capture technical metadata such as physical tables, fields, indexes, and constraints, as well as process metadata such as ETL (Extract-Transform-Load) task logic and data lineage, from the integrated ODS layer, business source database, or other systems. Tasks can be configured to be executed once or periodically to ensure the timeliness of metadata information.

[0097] ② Metadata Catalog Management: Provides a tree-structured directory management interface, allowing users to create multi-level directory structures based on business themes, departmental functions, and other dimensions. Collected metadata assets can be categorized and mounted to corresponding directories, and support the addition of business descriptions, responsible persons, and other business metadata, thereby forming a well-structured and information-rich enterprise data asset catalog, achieving data visualization and understandability.

[0098] 2) Data Governance Execution Submodule: Implement specific governance measures to ensure data consistency and standardization.

[0099] ① Data lineage analysis: Based on ETL task logs, SQL scripts, and table relationships collected from metadata, it automatically analyzes and constructs a complete data flow diagram from the original production system through multiple processing and integration processes, ultimately flowing to reports or application services. It supports interactive graphical display of upstream and downstream dependencies at the table and field levels, and can switch to a list view to view detailed lineage information for precise impact analysis and problem tracing. Impact analysis includes assessing the impact of source data changes on downstream processes, while problem tracing includes pinpointing the stages where data quality issues arise.

[0100] ② Master Data Management: Provides the ability to manage the entire lifecycle of core business entities, i.e., master data, for enterprises. Users can define master data models here, including attributes, validation rules, etc. The system supports full lifecycle management of master data, including CRUD operations, version control, and approval and release. Furthermore, by configuring distribution tasks, published standard master data can be synchronized to various downstream business systems, ensuring the consistency and authority of core data across systems.

[0101] ③ Data Tagging Management: Supports the construction of a flexible enterprise-level data tagging system for multi-dimensional and dynamic data classification and profiling. Users can define tag types and tag entities, and automatically or semi-automatically tag entities by configuring rules or connecting to machine learning models. The system provides a tag configuration interface and ad-hoc query functionality based on tag combinations, enabling refined data management and flexible retrieval. The tag type refers to the classification dimension, such as "customer value level" or "risk level"; the tag entity refers to the tagged object, such as "customer" or "product"; and the configuration rules include SQL rules and expressions.

[0102] 3) Data Quality Closed-Loop Management Submodule: Dedicated to measuring and improving the intrinsic quality of data, forming a management closed loop.

[0103] ① Quality Rule Management: Provides a centralized definition and repository for quality rules. It supports creating rule directories and defining specific rules within those directories. Rules are divided into two main categories: built-in rules, covering multiple dimensions such as completeness (not empty), uniqueness, accuracy (value range, regular expressions, logical consistency, etc.), and timeliness, which can be configured through graphical point-and-click functionality; and custom rules, allowing users to write Python scripts to implement more complex and personalized quality verification logic.

[0104] ② Quality Inspection and Evaluation: Allows users to create quality inspection tasks for specific data assets, flexibly combining and associating the aforementioned quality rules. Tasks can be scheduled and executed automatically, scanning the target data and producing two core results: a quality inspection log and a quality defect report. The quality inspection log records summary information such as the overall pass rate and rule triggering status for each inspection; the quality defect report lists all specific data records that violate the rules and their error values, supporting direct online correction.

[0105] ③ Intelligent quality enhancement: This sub-module deeply integrates AI capabilities to achieve an intelligent leap.

[0106] During the rule configuration phase, users can request AI assistance: the AI ​​large model intelligent engine module receives the structure of the target table, sample data, and business description, and automatically recommends potential quality risk points and corresponding rules through analysis. For example, it can identify that the "amount" field may have negative value anomalies, recommend the rule of "value range lower limit is 0", and generate the corresponding configuration code.

[0107] During the results analysis phase, the system can utilize an AI engine to perform in-depth processing of the "quality defect report": the AI ​​automatically clusters and performs attribution analysis on massive defect records, correlates data lineage and business metadata, and ultimately generates an intelligent quality assessment report. This report not only summarizes defect statistics but also describes the main quality issues in natural language, infers their business roots, assesses their potential impact on key business processes, and provides specific remediation priority suggestions, transforming technical quality inspection results into insights that directly support business decisions.

[0108] (3) AI Large Model Intelligent Engine Module

[0109] like Figure 4 As shown, this module is the soul and brain of this invention's "intelligent governance" mechanism. It is an independently deployed intelligent computing center that provides services externally via API. Its core lies in utilizing a large language model finely tuned with domain knowledge to inject three core AI capabilities into all the aforementioned functional modules:

[0110] 1) Intelligent Recommendation and Automated Configuration Generation: This is designed for configuration scenarios requiring specialized knowledge, such as data integration and access modules (identification rule selection) and core data processing and governance modules (quality rules, cleaning rules, data standard mapping configuration). When a user initiates a request, the engine receives relevant context such as data schema, samples, and business objective descriptions. Through model reasoning, it outputs the optimal configuration recommendation, such as "Recommend using the 'generalized simulation' identification rule for this 'customer name' field," or directly generates immediately executable task configuration code snippets, such as a Python script to check "date format consistency," greatly reducing the configuration threshold and improving efficiency and accuracy.

[0111] 2) Intelligent Report Generation and In-Depth Interpretation: Primarily serves the "Data Quality Closed-Loop Management Submodule" and "Data Lineage Analysis" functions within the core data processing and governance module. The engine receives structured input results such as quality defect lists and lineage graphs, combines them with rich business metadata such as asset catalogs and business terminology tables, and uses natural language generation and logical reasoning technologies to automatically generate structured and narrative analysis reports.

[0112] 3) Intelligent Interaction and Context-Aware Guidance: Serving as the driving engine for user interaction and intelligent assistance modules, it can understand user questions posed in natural language, such as "How do I configure an incremental synchronization task from MySQL to ODS?". It retrieves and organizes information from the system knowledge base and operation manual, providing clear step-by-step answers. Furthermore, based on the user's current system interface and ongoing operations, it proactively provides scenario-based and personalized operation guidance, achieving true intelligent assistance. For example, when creating quality rules, it prompts, "The table you currently select is often associated with the 'Customer Satisfaction Analysis' report; it is recommended to additionally configure a 'Customer ID' uniqueness rule."

[0113] (4) Data Assets and Services Module

[0114] This module is geared towards data consumers and aims to encapsulate, securely publish, and conveniently deliver governed, high-quality data assets, completing the "last mile" of data value realization.

[0115] 1) Data Asset Catalog: An asset portal for business users is built based on the metadata management submodule. All governed data tables, APIs, tag systems, and other assets are categorized and displayed in a more user-friendly way. Each asset comes with a detailed "asset card," including a business description, quality score, lineage, data sample, service level agreement (SLA), and application / usage entry point, making it easier for users to discover, understand, and evaluate data assets.

[0116] 2) Data as a Service: Supports encapsulating data query and processing logic into standardized, reusable API services. Users can configure the API's SQL query logic, input parameters, output format, etc., through the interface and conduct simulation tests. Once published, the API is included in a unified API catalog for lifecycle management, supporting API authorization management (associated with application AppKey), traffic monitoring, call log auditing, and online / offline control to ensure the security, stability, and controllability of data services.

[0117] 3) Service Operations and Security Gateway: Construct a secure access and control system for data assets. This includes: API permission management, assigning access keys (AppKey / AppSecret) to different applications (Apps) and setting access quotas and validity periods; service approval processes, where users submit data usage or API call applications through the "Asset View" or "My Applications" functions, and the process automatically flows to the relevant responsible persons for online service approval; and unified service monitoring, recording detailed logs of all API calls (API logs), including caller, interface, response time, status code, etc., for auditing, billing, and service quality analysis.

[0118] (5) User interaction and intelligent assistance module

[0119] This module serves as the unified front-end interface for the system's interaction with various user groups, such as... Figure 5 As shown. Besides providing traditional interactive components such as menu navigation, form filling, and list display, its core innovation lies in the comprehensive and deep integration of intelligent assistance capabilities:

[0120] 1) Traditional Functional Interface: Provides a clear and responsive web interface, navigating to major functional areas such as data integration, data governance, data quality, data services, and system management via a left-hand menu tree. Each functional area is designed according to a standard interaction pattern of "list browsing, search filtering, adding, editing, executing, and viewing details," ensuring efficient operation for professional users.

[0121] 2) Embedded Intelligent Assistance System: This is the key innovation of this module. A smart assistant dialogue pop-up is permanently embedded in the system interface. Users can ask questions in natural language at any time, such as: "Can you check if there are any recent problems with the data quality of the 'User Profile Table'?" or "What are the steps to create a new data anonymization task?". The smart assistant is driven by the AI ​​large-scale model intelligent engine module, which can not only understand questions, but also provide accurate answers, operation links, or even direct jumps based on real-time system status such as the current user role and page context. In addition, on key task configuration pages such as "Add Quality Rule", the system will dynamically sense user operations and provide an intelligent guidance panel in the sidebar or below the form, providing prompts for the current step, suggestions for avoiding common errors, or links to relevant best practice documents, achieving "immersive" teaching and assistance.

[0122] 3) Contextualized Intelligent Guidance and Prompts: When users are performing critical and complex tasks, such as creating a data development job involving multi-table joins and transformations, the system can sense the user's current step and provide intelligent guidance dynamically and non-intrusively. This may include: highlighting key configuration items for the current step, displaying common configuration suggestions or best practice tips for the step, automatically verifying the rationality of entered parameters and issuing warnings, and even predicting and recommending the next most likely operation based on the user's configured content, significantly reducing the learning curve and error rate of complex operations.

[0123] (6) System Management and Support Module

[0124] This module provides essential infrastructure, management tools, and operational support for the entire data governance system, ensuring that the system can operate securely, stably, and continuously.

[0125] 1) Basic Resource Management: Provides multi-level organizational structure management for institutions, positions, and users, enabling the mapping of users to entity organizations. Role-Based Access Control (RBAC) allows for fine-grained configuration of menus accessible to different user roles, operable buttons, and the range of data they can view (data permissions) through role management and menu management functions.

[0126] 2) System Configuration and Maintenance:

[0127] ① Data dictionary management: Maintain the standard code set of enumeration values, status codes and other data used globally by the system to ensure the consistency of similar data definitions in various modules.

[0128] ② Scheduled Task Center: Unifies the management of all jobs that need to be executed periodically, such as data integration, quality inspection, and metadata collection, and provides visualized CRON expression configuration, task status monitoring, and execution log viewing.

[0129] ③ File and Attachment Management: Unified storage and management of user-uploaded documents, images, and other attachment resources.

[0130] ④ Global monitoring dashboard (Operations and Maintenance Center): Centrally displays the running status, resource consumption, success / failure records of all batch processing jobs and stream processing jobs (such as Flink tasks), and provides operation and maintenance operation entry points such as job start and stop, log viewing, and alarm configuration, realizing transparent supervision of system operation status.

[0131] Through the collaborative work of the above six modules, this invention system constructs a complete data governance solution covering the entire data lifecycle from "access to governance to service to operation," with AI intelligence permeating the entire process. It not only inherits all the professional capabilities of mature data governance tools but also, through the deep empowerment of an AI large-scale model engine, achieves intelligent governance processes, insightful governance results, and universal system usability, effectively solving the various technical challenges described in the background section.

[0132] like Figure 2 As shown in the illustration, this invention provides a data governance method based on an AI large-scale model. The core of this method does not aim to describe a linear, mechanical data processing step. Its inventive point lies in constructing a collaborative computing paradigm that uses an AI large-scale model as the decision-making core and intelligent agents as the execution units, dynamically orchestrating and intelligently executing tasks throughout the entire data governance lifecycle. This process automatically decomposes, maps, and drives the collaborative work of various underlying functional modules through AI's understanding, planning, and scheduling of the user's high-level intentions or preset goals, thereby achieving automation, adaptability, and intelligence in the governance process. The main steps of the algorithm are as follows:

[0133] (1) Step S10: Governance task intent reception and intelligent parsing.

[0134] The system received a data governance task request. This request may originate from:

[0135] 1) Users input task instructions in the user interaction and intelligent assistance module through natural language, such as: "Prepare customer data for the marketing analysis report to be launched next month, ensure data quality and assess its lineage impact."

[0136] 2) Predefined smart tasks triggered by users on a certain function interface, such as clicking the "One-Click Smart Governance" button on the data table details page.

[0137] 3) Periodic governance tasks triggered by the timer of the system management and support module according to preset strategies (such as early morning every day).

[0138] Upon receiving a request, the process does not directly execute fixed operations. Instead, it first invokes the task understanding and decomposition service of the AI ​​large-scale model intelligent engine module. This service performs deep semantic parsing of the request, identifying core governance objectives (such as "data preparation," "quality assurance," and "impact assessment"), constraints (such as "serving marketing reports" and "time window: next month"), and key data assets involved (such as "customer data"). It then parses and transforms ambiguous user intent or simple instructions into one or more structured, actionable sets of governance sub-objectives. For example, the aforementioned natural language instructions might be parsed as: "Sub-objective A: Locate and integrate source data tables related to 'customers'," "Sub-objective B: Perform a comprehensive quality check and improvement on the integrated customer data," and "Sub-objective C: Analyze the lineage between the improved data and the 'marketing analytics report' and generate an evaluation report."

[0139] Step S20: Governance context awareness and resource discovery.

[0140] To achieve the parsed sub-goals, the system needs to fully understand the governance context related to the task. To this end, the process automatically triggers a context-aware agent. This agent autonomously calls APIs from multiple modules within the system to collect and integrate the following key information:

[0141] 1) Data Asset Context: From the metadata management submodule, query all data tables, fields and their business meanings related to "customer".

[0142] 2) Technical Status Context: Explore the current status, access method, update frequency, and historical data quality overview of relevant data tables from the source data storage and management layer and data integration and access module.

[0143] 3) Business Relationship Context: Obtain the lineage link between "customer" related data and downstream "marketing analysis reports" from the data lineage analysis function.

[0144] 4) Rule knowledge context: Retrieve existing, reusable quality rules and desensitization rule templates from the quality rule management and intelligent data desensitization layer.

[0145] 5) Environment and Constraint Context: Consider the current system load, available computing resources, and compliance policy requirements.

[0146] The significance of this step lies in providing rich and accurate factual evidence for subsequent intelligent decision-making and planning, so that AI's decisions are based on the real system state, rather than on reasoning out of thin air.

[0147] Step S30: Dynamic task planning and agent collaborative orchestration.

[0148] With a full understanding of the context, the planning and orchestration service of the AI ​​large-scale model intelligent engine module is activated. Based on the set of sub-objectives parsed in step S10 and the context information collected in step S20, it dynamically generates an optimal "governance execution plan".

[0149] This plan is not a static script, but a dynamic workflow completed collaboratively by multiple specialized intelligent agents. AI will determine which types of agents are needed, such as a data exploration and integration agent, a quality governance agent, and a report generation agent, and will assign appropriate sub-tasks, execution parameters, triggering conditions, dependencies, and data transfer methods to each agent. For example, the planned outcome might be: "First, the 'data exploration and integration agent' confirms and executes the latest synchronization of customer data; upon completion, it triggers the 'quality governance agent' to recommend, configure, and execute quality checks and cleansing based on intelligent rules; finally, the 'report generation agent' summarizes all process logs and quality results to generate a final comprehensive evaluation report."

[0150] Step S40: Modular execution and closed-loop feedback driven by intelligent agents.

[0151] Each specialized intelligent agent is instantiated and begins execution according to the planned workflow. Each agent is essentially a software entity that encapsulates domain-specific knowledge such as data integration, quality inspection, and operational logic (calling corresponding system module APIs), and is supported by a large AI model for decision-making.

[0152] 1) Execution process: While executing business logic such as calling the fast data integration API, the intelligent agent will request the intelligent recommendation service of the AI ​​large model intelligent engine module to provide real-time suggestions at key decision points such as selecting de-identification rules and recommending quality rules, or directly submit the structured intermediate results generated during execution, such as a list of quality defects, to the intelligent report generation service for real-time analysis and interpretation.

[0153] 2) Closed-loop feedback: The AI ​​rigorously monitors the execution status and results of every underlying API it calls. If a step fails or the output does not meet expectations (e.g., a low quality check pass rate), this information is sent back to the planning and orchestration service in real time. The AI ​​can then dynamically adjust subsequent plans accordingly, such as replanning remedial tasks, trying alternative execution paths, or deciding to report the anomaly to the user for manual intervention. This closed-loop mechanism of "perception-planning-execution-feedback" endows the entire process with strong fault tolerance and adaptability.

[0154] Step S50: Governance outcome aggregation and adaptive delivery.

[0155] When all agents complete their tasks, or the process reaches the predetermined termination conditions, the results aggregation phase begins. A results aggregation agent collects all key deliverables generated throughout the entire intelligent governance process, including: executed task logs, newly created or modified governance rule configurations, generated data quality scores, intelligent evaluation reports, etc.

[0156] Ultimately, the process will deliver the results in the most appropriate way based on the initial source of the task request:

[0157] 1) For requests from the intelligent assistant, the results will be presented to the user in the dialogue interface in the form of a structured natural language summary, along with a link to a detailed report.

[0158] 2) When triggered by a function button, the system will automatically navigate the user to a comprehensive results dashboard page, which will provide a panoramic view of the governance process and results.

[0159] 3) For tasks that are triggered on a schedule, the results will be recorded and may trigger notifications or be included in the daily operation and maintenance report.

[0160] Step S60: Knowledge accumulation and model iterative optimization (continuous learning phase).

[0161] To enhance the system's long-term intelligence, this algorithm incorporates an offline, continuous learning component. The system stores the execution records of each complete process, contextual information, AI decision points, and final user feedback (such as adoption / modification of the generated report) as training data. This data can be periodically used to fine-tune or optimize the core models in the AI's large-scale intelligent engine module through reinforcement learning. This allows the AI's intent understanding, task planning, rule recommendation, and report generation capabilities to continuously evolve with accumulated practical experience, more accurately matching the specific business scenarios and user habits of enterprises.

[0162] In summary, the method proposed in this invention is an advanced approach driven by a large AI model and achieved through multi-agent collaboration to dynamically orchestrate and intelligently execute data governance tasks. It transcends the rigid, linear workflow of traditional methods, transforming data integration, desensitization, governance, and quality evaluation into flexible services that can be dynamically scheduled and intelligently decided by AI. This systematically solves the core problems of low intelligence and poor adaptability mentioned in the background art, constituting the key methodological innovation of this invention.

[0163] This embodiment uses the data governance needs of a health technology company as a case study to elaborate on how the AI-based big data model-based data governance system and algorithm proposed in this invention can intelligently govern the core data asset, the "user detection data table," throughout its entire lifecycle, and ultimately generate an intelligent quality assessment report that can directly support business decisions. This case study will fully demonstrate the entire process from raw data access to intelligent service delivery, highlighting the intelligent leap forward of this invention compared to traditional data governance tools.

[0164] 1. Scenario and Data Background

[0165] A company operates a digital health management platform, where "user monitoring data" is a key data asset supporting core businesses such as chronic disease management and sub-health monitoring. This data is scattered across multiple sources:

[0166] The core business database (db_biz_core) stores user-submitted testing applications and results; the core table is src_user_exam.

[0167] Device Management Database (db_device_mgr): Records metadata and status of various health monitoring devices. The src_device_info table contains information such as device model and accuracy.

[0168] The user center database (db_user_center) stores basic user information. The table src_user_profile contains user ID, demographic information, etc.

[0169] Governance Objective: Data administrator Xiao Wang needs to ensure the high quality and credibility of "user detection data" and regularly generate "User Detection Data Quality Intelligent Assessment Report" to assess the status of data assets, guide business optimization, and meet internal and external audit requirements.

[0170] 2. Intelligent governance steps throughout the entire process based on the present invention

[0171] Step 1: Data Integration and Access

[0172] After logging into the system, Xiao Wang directly entered the following through the intelligent assistant floating window of the user interaction and intelligent assistance module: "I need to connect to the user detection table in the production database to conduct a comprehensive data quality governance."

[0173] (1) Intent understanding and intelligent guidance

[0174] The AI ​​large-scale model intelligent engine module parses instructions in real time, identifies keywords such as "user detection table" and "data quality governance," and automatically associates them with the possible data source table user_exam by combining them with the system's metadata knowledge base.

[0175] The intelligent assistant responded immediately: "We have detected that you may need to access the user_exam table in the db_health database. We have navigated you to the 'Data Integration and Access' module. Do you want to be guided through the configuration process step by step?" Xiao Wang clicked "Yes".

[0176] Step 2: Data Access and Security Desensitization Under Intelligent Guidance

[0177] After logging into the system, Xiao Wang didn't need to remember complicated paths. He directly entered natural language commands in the intelligent assistant floating window of the user interaction and intelligent assistance module: "I need to connect to the business data related to user detection for subsequent quality assessment and report generation."

[0178] (2) Intent understanding and intelligent navigation:

[0179] The AI-powered big data engine module analyzes commands in real time, identifying the core intent as "data access" and related to the "user detection" business. Combining this with the system's registered data source directory and metadata knowledge base, the AI ​​intelligently infers that data sources such as db_biz_core, db_device_mgr, and db_user_center may be involved.

[0180] The intelligent assistant responded immediately: "We understand you need to access 'user detection' related data. Three potential relevant data sources have been identified. I will guide you through the connection configuration. First, I will navigate you to the 'Data Source Management' interface." The system interface then automatically redirected, launching the intelligent guide panel in the sidebar.

[0181] Multi-source connectivity configuration and AI-assisted desensitization:

[0182] Following the onboarding prompts, Xiao Wang began configuring the connection to the db_biz_core data source. When browsing the field list of the src_user_exam table, he was unsure about the anonymization strategies for "User ID Card Number" (assuming the field name is user_idcard) and "Detailed Data" (which may contain sensitive health indicators).

[0183] AI-powered intelligent de-identification recommendation: Xiao Wang selects the `user_idcard` field and clicks "AI Recommend De-identification Rules". The AI ​​large-scale model intelligent engine is triggered, providing recommendations based on the following analysis:

[0184] Field semantics and sample identification: This field is identified as a personal ID number.

[0185] Compliance matching: Based on laws and regulations such as the Personal Information Protection Law, it is determined to be highly sensitive personal identity information.

[0186] Business scenario considerations: Considering that subsequent statistical analysis may be required based on region (first few digits of ID card), but it is not necessary to be accurate down to the individual level.

[0187] Recommendation result: The AI ​​highlighted the following recommendation on the interface: "To balance privacy protection and data analysis value, it is recommended to use a 'mask (keeping the first 6 digits and the last 4 digits)' rule for the 'user ID number' field, for example, '1101011234'." Xiao Wang adopted the recommendation with one click.

[0188] Similarly, for the `detailed_data` field, AI analysis suggests it may contain sensitive health information, but business analysis requires numerical trends. Therefore, the "numerical perturbation (adding small random noise within a reasonable range)" rule is recommended to protect privacy while maintaining statistical characteristics. Xiao Wang confirmed the adoption of the AI's explanation.

[0189] (3) Automated data extraction and lake entry:

[0190] Guided by the AI ​​assistant, Xiao Wang proceeded to create a "rapid data integration" task. Based on his initial intent, the AI ​​pre-populated the task name as "User Detection Core Data Access" and recommended a scheduling cycle of "Daily Full and Incremental Hybrid Synchronization".

[0191] Based on the table names and field comments, the AI ​​intelligently recommended integrating three tables: src_user_exam (core fact table), src_device_info (dimension table), and src_user_profile (dimension table), and suggested using exam_id (detection serial number) and user_id (user ID) as key related fields.

[0192] After task configuration, the process executes automatically. Source data (after anonymization) is stably synchronized to the system's source data storage and management layer (ODS), generating ods_user_exam, ods_device_info, and ods_user_profile respectively. Complete end-to-end access logs are recorded. Step Two: Intelligent Data Governance and Quality Assurance

[0193] After the data was fed into the lake, Xiao Wang continued to issue instructions through the smart assistant: "Data has been connected. Please perform comprehensive quality management on the ods_user_exam table."

[0194] (1) Automatic metadata collection and asset cataloging:

[0195] The system backend automatically triggered the metadata management submodule's data collection task to obtain the technical metadata of the newly added tables. Subsequently, AI suggested mounting the ods_user_exam table to the asset directory under "Business Themes > Health Management > User Detection" and automatically generated a preliminary business description: "Records core factual data of user health detection, including detection items, raw data, analysis results, etc." Xiao Wang added information such as the data manager, forming a complete asset card.

[0196] (2) AI-driven intelligent configuration of quality rules (core innovation):

[0197] Xiao Wang navigated to the "Data Quality" module and created a quality check task for the ods_user_exam table. Faced with numerous fields and complex rule options, he clicked the "AI Intelligent Recommendation Rules" button.

[0198] 1) In-depth analysis of the AI ​​engine:

[0199] ① Understanding Field Semantics and Structure: Identify that examid is the primary key, userid is the user foreign key, examinee is the category field, data and result are text-based results, and created is the timestamp.

[0200] ② Data Sample Scanning and Pattern Discovery: After quickly scanning the data samples, the following was discovered:

[0201] The examcate (test item) field contains an excessively long string (exceeding the defined length).

[0202] The created (time) field is in standard DATETIME format, but there are a very small number of future timestamp records (such as due to server time synchronization issues).

[0203] There may be a few null values ​​for userid, but the business logic may allow anonymous detection, so it needs to be judged in conjunction with the rules.

[0204] The data field (collected data) is MEDIUMTEXT, and its JSON / XML format specification needs to be considered (if it is agreed to be structured text).

[0205] ③ Matching industry rules with best practices: In accordance with the "Medical and Health Data Management Standard", it is determined that the test serial number should be unique, the time should be reasonable, and the key classification information should be standardized.

[0206] ④ AI-powered rule-based recommendation list generation: The AI ​​clearly displays the recommendation rules on the interactive interface.

[0207] target field Recommended rule types Rule logic / parameters Reasons for recommendation and explanation examid Uniqueness check The field value is unique across the entire table. The serial number is the core of the data primary key and association, and its uniqueness must be guaranteed to avoid data duplication and association errors. userid Not empty check Field values ​​are not allowed to be NULL. User identification is crucial for linking user profiles and conducting personalized analysis. If the business allows anonymous detection, it can be configured as "warning" instead of "blocking". examcate Standardization - Length Check The string length is ≤ 50 characters. The names of test items should be concise and standardized. Excessively long descriptions in samples may affect statistical classification and front-end display. It is recommended to promote standardization at the source. examcate Normative - Range Reference The value should be within the standard test item dictionary (e.g., blood glucose, blood pressure, electrocardiogram, blood lipids). This improves data consistency, facilitating subsequent indicator calculations and multidimensional analysis. The system can be validated by referencing a pre-defined data dictionary. source Not empty check Field values ​​are not allowed to be NULL. Source equipment information is crucial for assessing the reliability of data collection and tracing equipment-related issues. created Accuracy-Reasonability Verification Date and time ≤ current system time, and ≥ the system's preset business start time (such as the platform launch date). The detection time should not be a future time, nor should it be earlier than the reasonable starting point of the business, to ensure the correctness of the business time logic. data Standardization - Format Validation The content is a valid JSON format string (if this format is agreed upon). As a semi-structured data container, the format of the data field directly affects the success or failure of subsequent parsing and analysis. (Table level) Custom logic rules The same user (userid) should not have more than N test records for the same test item (examcate) within 24 hours (e.g., N=5). Rules for preventing duplicate and fraudulent transactions based on business common sense are used to avoid interference from abnormal data. AI can provide Python script templates.

[0208] (3) User interaction, adoption and enhancement:

[0209] Xiao Wang reviews the AI ​​recommendation list.

[0210] He accepted the following: He considered the uniqueness of examid, the non-empty userid, the reasonableness of created, and the data format validation to be fundamental and necessary, and directly selected "accept". For the non-empty userid rule, he accepted the AI's suggestion and set the rule strength to "warning".

[0211] Adjustment: Regarding the value range rules for "examcate", he found that the standard dictionary lacked new items such as "gene testing". He adopted the rule, but modified the rule pattern to "trigger a warning and log values ​​not in the dictionary for future dictionary expansion".

[0212] Additional information: Based on his business experience, Xiao Wang added an extra custom rule: "If the result field contains keywords such as 'critical' or 'serious anomaly,' then the examinee field and the created field must not be empty, and the created time must be within the last 7 days." He quickly completed the configuration using an AI-generated script framework.

[0213] (4) Execution of inspection and intelligent closed-loop repair:

[0214] Once the rules are configured, the quality check task is executed immediately. The system generates preliminary results: 500,000 records were checked, with an overall pass rate of 99.1%. The main defects are concentrated in excessively long `examcate` entries (approximately 0.5% of records) and a small number of `created` entries with future time points.

[0215] Xiao Wang clicked "AI Generate Repair Suggestions." The AI ​​engine then performed a deep analysis of the defect record:

[0216] 1) Clustering and Attribution: "There are 2580 excessively long questions in the examcate, 80% of which come from 'Device Model A'. Its firmware version merges 'Detection Method' and 'Item' in the output. It is recommended to split the fields at the data integration layer or through a cleaning job."

[0217] 2) Impact Assessment: "There are 12 future timestamps created, involving 2 test devices. Although the number is small, it can cause serious bias in time series analysis. Recommendation: Correct the timestamps to the correct ones in batches based on the actual received logs."

[0218] 3) Repair Script Generation: The AI ​​automatically generated executable SQL cleaning scripts and Python data processing code snippets to address the two issues mentioned above. After reviewing them, Xiao Wang added the cleaning scripts to the governance task flow and executed them with a single click, achieving an intelligent closed loop from problem discovery and root cause analysis to repair execution.

[0219] Step 3: Generation and Interpretation of AI-Driven Intelligent Quality Assessment Reports

[0220] After the core table governance and repair were completed, Xiao Wang issued the final instruction to the intelligent assistant: "Now, based on the latest data, generate a comprehensive intelligent quality assessment report."

[0221] (1) Agent-driven panoramic evaluation:

[0222] The AI ​​task parsing engine interprets this instruction as the goal of "generating an ods_user_exam intelligent quality assessment report". The context-aware agent automatically collects: the final quality inspection results, repair logs, table lineage information (related to ods_device_info, ods_user_profile), and business importance tags in the asset catalog (marked as "core business table").

[0223] The AI ​​planning service schedules a "report generation agent." This agent invokes the deep analysis service of the AI ​​large-scale model intelligent engine and inputs all structured data.

[0224] (2) In-depth analysis and narrative report generation:

[0225] The AI ​​report generation service, based on input, performs the following operations to generate the "Intelligent Quality Assessment Report for User Detection Data Assets (ods_user_exam)," such as... Figure 6 As shown.

[0226] (3) Intelligent delivery and knowledge accumulation:

[0227] 1) Intelligent Delivery: The report is automatically published to the asset catalog of the Data Assets and Services module as an important subsidiary asset in the ods_user_exam table, and a subscribed API is generated. Simultaneously, the intelligent assistant pushes a notification to Xiao Wang: "The intelligent quality report for your requested user detection data has been generated. Overall score: 99.8 (Grade A). Core conclusion: Historical compliance issues have been closed, data assets are in good health, and it is recommended to monitor the status of associated devices. Report details and API are ready."

[0228] 2) Continuous learning: Throughout the entire governance process, the decision-making logic, contextual data, and final effect feedback of AI in stages such as de-identified recommendation, rule generation, root cause analysis, and report insight are all automatically encrypted and recorded by the system to form high-quality training data. This data is then fed back into the continuous learning pool of the AI ​​large model intelligent engine module for subsequent incremental fine-tuning and optimization of the model, enabling its intelligence level in the field of medical and health data governance to continuously evolve.

[0229] This embodiment takes the intelligent governance and quality assessment of the "user detection data table" as a clue, and fully and meticulously demonstrates how the system of the present invention starts from a natural business instruction, and through the deep empowerment of the AI ​​big model intelligent engine module, drives the collaborative work of modules such as data integration and access, core data processing and governance, user interaction and intelligent assistance, etc., throughout the entire process of secure data access, intelligent desensitization, automatic configuration and recommendation of quality rules, intelligent defect repair, and finally the generation of an intelligent assessment report with in-depth business interpretation.

[0230] Compared to traditional data governance tools, the key advantages of this invention in this embodiment are:

[0231] Intelligent configuration: The complex selection of desensitization strategies and configuration of quality rules are transformed from "manual work" that relies heavily on expert experience into an "intelligent assistance" mode that is recommended by AI analysis and adopted by users with one click, greatly reducing the technical threshold.

[0232] Closed-loop governance: It not only discovers problems (such as excessively long fields), but also uses AI to deeply analyze and locate the technical root cause and business source of the problem, and provides directly executable remediation solutions (such as cleaning scripts), realizing an intelligent closed loop of "discovery-analysis-remediation".

[0233] Results Insight Transformation: Transform dry quality score lists and defect details into narrative-driven intelligent reports that include scores, root causes, business impact, associated risks, and specific action recommendations. This enables business and management personnel to directly understand and apply the results of technology governance, unlocking the business value of data governance.

[0234] Humanized Interaction: The entire process is guided by a natural language assistant and scenario-based prompts, enabling data administrators like Xiao Wang to drive complex technical processes using business language, significantly improving work efficiency and system adoption.

[0235] This embodiment fully verifies that the data governance paradigm shift from "tool-assisted" to "intelligent-driven" achieved by the present invention through the "AI large-scale model intelligent engine" and the "algorithm process of intelligent agent collaboration" effectively solves the core defects proposed in the background technology, such as the reliance on manual configuration of governance rules, the lack of in-depth insight in governance reports, and the complexity and difficulty in using the system operation.

[0236] Example 1: Intelligent Planning and Execution Closed-Loop Implementation Based on Governance Intent Understanding

[0237] In this embodiment, the data governance system is deployed in an enterprise private cloud environment. When business personnel submit a governance task to "improve customer data quality" through the system, the system first sends the task request to the AI ​​big data engine. The engine performs semantic parsing on the input natural language task, identifies the governance object as customer-related data, and determines that the governance goal is to improve integrity, consistency, and accuracy. It then automatically breaks down this governance goal into multiple sub-tasks, including field integrity detection, duplicate record identification, and outlier analysis. Subsequently, the system generates a structured governance execution plan based on the breakdown results and distributes it to the data governance execution environment.

[0238] During execution, the system continuously receives intermediate results generated by the governance process, such as the hit rate of quality rules and the distribution of anomalies. When persistent quality anomalies are detected in some fields, the AI ​​large-scale model intelligent engine automatically adjusts subsequent execution steps, supplementing with stricter verification strategies and reordering the execution sequence. This process, without manual reconfiguration, forms a complete closed loop from task understanding and governance planning to execution feedback and optimization, demonstrating that the system possesses stable and repeatable data governance capabilities.

[0239] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0240] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A data governance system based on AI large-scale models, characterized in that, The system includes a data governance execution environment and an AI big model intelligent engine, which are connected through a standardized interface. The AI ​​big model intelligent engine is used to perform semantic understanding and target decomposition of the governance task after receiving a data governance task request. Based on the decomposition results, it generates a governance execution plan that matches the data governance execution environment. During the governance execution process, it dynamically adjusts the governance execution plan based on the execution feedback, thereby forming a data governance closed-loop mechanism that runs through task understanding, planning, execution and feedback optimization.

2. The system as described in claim 1, characterized in that, When generating a governance execution plan, the AI ​​big model intelligent engine automatically determines the order of governance steps, the relationship between governance rules and execution dependencies based on governance task objectives, data structure information and historical governance knowledge, so as to form a structured governance process that can be directly driven to execute.

3. The system as described in claim 1, characterized in that, The dynamic adjustment includes: During the governance implementation process When changes in data quality, rule execution anomalies, or resource status changes are detected The AI-powered large-scale model intelligent engine re-plans governance steps that have not been executed or are currently being executed. And send the updated execution instructions to the data governance execution environment.

4. A data governance system based on AI large-scale models, characterized in that, The system includes multiple specialized intelligent agents for different governance functions and a collaborative scheduling mechanism. The specialized intelligent agents include at least a data discovery agent, a rule reasoning agent, and a result aggregation agent. Under the control of the AI ​​big model intelligent engine, the collaborative scheduling mechanism instantiates, schedules, and orchestrates the specialized intelligent agents based on the context information of the governance task to achieve collaborative execution across governance stages.

5. The system as described in claim 4, characterized in that, The context information includes at least data asset context, technical status context, business relationship context, and rule knowledge context. The context information is provided by the corresponding system module through an interface and is continuously updated during the execution of the governance process.

6. The system as described in claim 4, characterized in that, The result aggregation agent is used to uniformly summarize the governance results, execution logs, and quality assessment results generated by each professional agent when the governance process ends or when predetermined termination conditions are met, and to generate delivery results that match the source of the governance task.

7. A data governance system based on AI large-scale models, characterized in that, The system constructs a governance knowledge accumulation mechanism to structurally store the task intent, contextual information, governance decision-making process, execution results, and user feedback in each data governance process. This governance knowledge is then used as training samples to periodically optimize the AI ​​large-scale model intelligent engine, thereby improving the automation level and decision accuracy of subsequent governance tasks.

8. The system as described in claim 7, characterized in that, The governance knowledge includes governance rule selection paths, task breakdown results, execution order adjustment records, and result evaluation information.

9. The system as described in claim 7, characterized in that, The periodic optimization includes updating the parameters or prompting strategies of the AI ​​big model based on historical governance success rates and user feedback results, thereby reducing the need for manual intervention in subsequent governance tasks.

10. The system as described in claim 7, characterized in that, The governance knowledge accumulation mechanism supports the classification and management of governance knowledge according to business themes or data asset categories, so as to realize the reuse of governance experience for different business scenarios.