Data lake treatment system and method based on digital twinborn and dynamic mapping knowledge domain
The data lake governance system, which utilizes digital twins and dynamic knowledge graphs, has solved the problems of data collection being disconnected from business operations and insufficient quality auditing. It has achieved automated and intelligent data governance, improved data accuracy and response speed, and met personalized analysis needs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHANJIANG BRANCH OF CHINA NATIONAL OFFSHORE OIL CORP
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Existing enterprise data lake construction suffers from problems such as data collection being disconnected from business operations, outdated quality assessment methods, fragmented toolchains, and long data preparation cycles, failing to meet the data governance needs driven by business processes.
A data lake governance system based on digital twins and dynamic knowledge graphs is adopted, including a business digital twin module, an intelligent collection triggering and verification module, a data access and dynamic quality governance engine, and a low-code/no-code self-service analysis platform, to achieve automation and intelligence in data collection and quality auditing.
It enables proactive, accurate, and timely full lifecycle data governance, improving data accuracy and business relevance, reducing operation and maintenance costs and error rates, and quickly responding to personalized analysis needs.
Smart Images

Figure CN121958249A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of enterprise information management technology, and in particular relates to a data lake governance system and method based on digital twins and dynamic knowledge graphs. Background Technology
[0002] A data lake is a centralized storage system capable of holding structured and unstructured data of any size. Unlike data warehouses, which are designed specifically for particular analytical tasks, data lakes preserve the raw form of data until it is processed, thus supporting a wide range of data analysis activities. Data lakes can be deployed on cloud platforms or on-premises data centers, and their flexibility and scalability are well-suited to modern big data needs.
[0003] Currently, enterprise data lake construction generally faces the dilemma of "emphasizing storage while neglecting governance." Existing technical solutions are mostly combinations of isolated tools, which have the following inherent defects: 1. Data collection is disconnected from business operations: Data collection relies on manual initiative and is not mandatory to project schedules, resulting in delays and omissions in data entry into the data lake.
[0004] 2. Outdated quality assessment methods: Quality checks are mostly based on static rules (such as NOT NULL and format), which cannot identify complex business logic errors and semantic-level "inaccuracy" issues.
[0005] 3. Fragmented toolchain: The integration between data acquisition tools, data lakes, and analysis applications is low, and the synchronization strategy is rigid, often requiring manual intervention.
[0006] 4. Long data preparation cycle: Traditional ETL processes cannot meet the agile and personalized analysis needs of business users such as researchers, and the release of data value is slow.
[0007] While existing technologies involve data quality monitoring, they do not address the issue of business process-driven data collection; although knowledge graphs are mentioned, they are not applied to dynamic data auditing and intelligent generation of data retrieval lists.
[0008] Therefore, there is an urgent need for an integrated governance solution that can deeply integrate business flows and data flows. Summary of the Invention
[0009] The problem this invention aims to solve is to provide a data lake governance system and method based on digital twins and dynamic knowledge graphs. This method addresses the technical shortcomings of existing data lake management, such as the disconnect between business and data, superficial quality audits, and slow data application, and realizes a proactive, accurate, timely, and agile full lifecycle data governance system.
[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a data lake governance system based on digital twins and dynamic knowledge graphs, characterized in that it includes: The business digital twin module is used to build and maintain digital twin models of business projects, and the digital twin models synchronize the status of key nodes of the business projects in real time. The intelligent acquisition triggering and verification module is used to automatically trigger and verify acquisition tasks based on the model status; The data access and dynamic quality governance engine is used to receive data from source collection tools and use a pre-built dynamic knowledge graph to perform semantic-level correlation and compliance audits on the received data. A low-code / no-code self-service analytics platform is used to respond to users' analytical operation intentions and automatically generate and recommend a list of data to be retrieved based on the dynamic knowledge graph to complete the analysis.
[0011] Furthermore, the data access and dynamic quality governance engine includes, The adaptive synchronizer is configured to receive data streams actively pushed by the source tools via a standardized API; The dynamic knowledge graph construction unit is used to automatically build and update domain knowledge graphs describing data entities, business rules and their relationships from enterprise data standard documents, business system metadata and historical data question databases; A semantic-level quality auditor is used to perform semantic-level correlation and compliance checks on the data entering the lake using the domain knowledge graph.
[0012] Furthermore, the low-code / no-code self-service analysis platform includes a visual orchestration environment, an integrated coding environment, and an intelligent data retrieval list generator. The intelligent data retrieval list generator is used to query the dynamic knowledge graph based on the user's analysis intent on the platform, automatically reason, and generate a standardized data retrieval list.
[0013] Furthermore, this invention also provides a data lake governance method based on digital twins and dynamic knowledge graphs, utilizing the aforementioned intelligent management and optimization system based on exploration big data, including the following steps: S1: Construct a digital twin model of the business project, wherein the digital twin model synchronizes the status of key nodes of the business project in real time; S2: Monitor the node status changes of the digital twin model, and when the status meets the preset conditions, automatically generate and send a data acquisition task to the designated terminal; S3: Receive data from source collection tools and use a pre-built dynamic knowledge graph to perform semantic-level correlation and compliance audits on the received data; S4: In the self-service analysis platform, in response to the user's analysis operation intention, the platform automatically generates and recommends a list of data to be retrieved to complete the analysis based on the dynamic knowledge graph.
[0014] Furthermore, in S2, a dynamic verification rule is built into the data acquisition task to perform preliminary format and logic verification when the data is submitted, and the task completion status is fed back to the digital twin model.
[0015] Furthermore, in S3, information is extracted from enterprise data standard documents, business system metadata, and historical data problem databases. Through entity recognition and relationship extraction technologies, a graph structure describing data assets and their business rules is automatically constructed and continuously updated to build the dynamic knowledge graph.
[0016] Furthermore, in S3, based on the dynamic knowledge graph, entities and business rules associated with the audited data are queried, and cross-field and cross-table logical consistency checks and business threshold judgments are performed to conduct semantic-level correlation and compliance audits.
[0017] Furthermore, in step S4, the data retrieval list includes: the standard field names, formats, business meanings, and suggested sources for the required data.
[0018] Furthermore, the present invention also provides a computer device, including a memory, a processor, and an algorithm stored in the memory and executable on the processor, wherein the processor implements the above-described data processing method when executing a computer program.
[0019] Furthermore, the present invention also provides a computer-readable storage medium storing a computer algorithm, which, when executed by a processor, performs the above-described data processing.
[0020] The advantages and positive effects of this invention are: 1. This invention provides source protection and achieves both quality and efficiency. By using digital twin-driven data collection, it ensures the integrity and timeliness of data from the source.
[0021] 2. This invention features an intelligent deep audit function, which achieves semantic-level quality checks through a dynamic knowledge graph, significantly improving the accuracy of data and its relevance to business needs.
[0022] 3. This invention is highly automated, from data collection and triggering to data synchronization and quality inspection, greatly reducing manual intervention, lowering operation and maintenance costs and error rates.
[0023] 4. This invention features agile business empowerment, a low-code / no-code platform, and an intelligent data retrieval list, which greatly reduces the threshold for data use, quickly responds to personalized analysis needs, and directly empowers intelligent research and scientific decision-making. Attached Figure Description
[0024] Figure 1 This is an overall architecture framework diagram of an embodiment of the system of the present invention.
[0025] Figure 2 This is a flowchart of the intelligent acquisition triggering method based on digital twins in an embodiment of the present invention.
[0026] Figure 3 This is a flowchart of a data quality auditing method based on dynamic knowledge graphs, as described in an embodiment of the present invention.
[0027] In the picture: 101. Business Digital Twin Module; 102. Intelligent Collection Triggering and Verification Module; 103. Data Access and Dynamic Quality Governance Engine; 104. Low-code / No-code Self-service Analysis Platform. Detailed Implementation
[0028] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] The embodiments of the present invention will be further described below with reference to the accompanying drawings: A data lake governance approach based on digital twins and dynamic knowledge graphs includes the following steps: S1: Construct a digital twin model of the business project, which synchronizes the status of key nodes of the business project in real time.
[0030] S2: Monitor the node state changes of the digital twin model. When the state meets preset conditions, automatically generate and send a data acquisition task to the designated terminal. Build dynamic verification rules for the acquisition task, perform preliminary format and logic checks when data is submitted, and feed back the task completion status to the digital twin model.
[0031] S3: Receive data from source data collection tools and utilize a pre-built dynamic knowledge graph to perform semantic-level relevance and compliance audits on the received data. Information is extracted from enterprise data standard documents, business system metadata, and historical data issue databases. Through entity recognition and relationship extraction technologies, a graph structure describing data assets and their business rules is automatically constructed and continuously updated, forming the dynamic knowledge graph. Based on this dynamic knowledge graph, entities and business rules associated with the audited data are queried, and cross-field and cross-table logical consistency checks and business threshold judgments are performed to conduct semantic-level relevance and compliance audits.
[0032] S4: In the self-service analysis platform, in response to the user's analytical intent, the system automatically generates and recommends a data retrieval list based on the dynamic knowledge graph to complete the analysis. The data retrieval list includes: the standard field names, formats, business meanings, and suggested sources for the required data.
[0033] The present invention also provides a data lake governance system based on digital twins and dynamic knowledge graphs, which runs the above-mentioned intelligent management optimization method based on exploration big data, including a business digital twin module 101, an intelligent acquisition triggering and verification module 102, a data access and dynamic quality governance engine 103, and a low-code / no-code self-service analysis platform 104.
[0034] The Business Digital Twin Module 101 is used to build and maintain digital twin models of business projects. Specifically, it interfaces with external project management systems to construct digital twin models of project processes. This model dynamically reflects key milestones such as the project's "six nodes" and their status. It is the first time digital twin technology has been applied to data acquisition planning, achieving a precise mapping from business status to data acquisition tasks.
[0035] The intelligent data acquisition triggering and verification module 102 monitors changes in the node status of the digital twin model, automatically triggering and verifying data acquisition tasks based on the model's status. Specifically, when a node reaches a predetermined state, a data acquisition task is automatically generated and pushed to the designated responsible person. The task has a built-in dynamic verifier that performs basic verification immediately upon data submission. This automatic triggering and preliminary verification mechanism based on business process status transforms data acquisition from "post-event supplementation" to "in-process generation," fundamentally ensuring timeliness and completeness.
[0036] The data access and dynamic quality governance engine 103 includes an adaptive synchronizer, a dynamic knowledge graph construction unit, and a semantic-level quality auditor.
[0037] Specifically, the adaptive synchronizer provides standardized APIs and configurable connectors to support automatic, real-time push from source tools to the data lake.
[0038] Dynamic Knowledge Graph Construction Unit: Automatically constructs and updates a domain knowledge graph describing data entities, business rules, and their relationships from enterprise data standards, business specifications, and historical question databases.
[0039] Semantic-level quality auditor: Utilizing the aforementioned knowledge graph, it performs semantic-level correlation and compliance checks on the data entering the lake. For example, it checks whether the correlation parameters between "daily fluid production of oil wells" and "oil pressure" and "casing pressure" are within the reasonable range defined by the knowledge graph.
[0040] Problem closed-loop management unit: Automatically generates work orders for problems identified during audits, and includes them in a list for tracking and resolution.
[0041] Real-time semantic-level data quality auditing based on dynamic knowledge graphs surpasses traditional rule-based checks and can intelligently discover deep-seated, interconnected data errors.
[0042] The 104 Low-Code / No-Code Self-Service Analytics Platform includes a visual orchestration environment, an integrated coding environment, and an intelligent data retrieval list generator.
[0043] Specifically, the visual orchestration environment provides drag-and-drop components, enabling users to build data pipelines and analysis reports in a low-code manner.
[0044] Integrated coding environment: Seamlessly embedded in professional coding environments (such as Jupyter) to meet the complex processing needs of advanced users.
[0045] Intelligent Data Extraction List Generator: Based on the user's analytical intent on the platform (such as the selected business theme, dragged-in dimensions and metrics), it queries the dynamic knowledge graph, automatically infers and generates a standardized "data extraction list" to ensure the completion of the analysis. This list can guide the improvement of source data standards.
[0046] The intelligent data retrieval list generation based on analytical intent and knowledge graph forms an intelligent closed loop of "application-driven governance," solving the problem of difficulty in implementing standards.
[0047] The following is in conjunction with the appendix Figure 1-3 The present invention will be described in detail below: like Figure 1 As shown, this is the data lake governance system based on digital twins and dynamic knowledge graphs provided by the present invention.
[0048] Among them, the vertical data flow shows the complete path of data from the external system layer, through the core system layer of this invention, and finally into the data storage layer.
[0049] Horizontal business flow: This demonstrates the business closed loop from the business digital twin module 101 to the intelligent collection trigger module 102, then to the data access and dynamic quality governance engine 103 for assurance, and finally to the consumption of value through the low-code / no-code self-service analysis platform 104.
[0050] Core Interactions: The Business Digital Twin Module 101 <-> Intelligent Data Acquisition Trigger Module 102 demonstrates the strong coupling between business processes and data acquisition tasks.
[0051] The Data Access and Dynamic Quality Governance Engine 103 <-> Low-Code / No-Code Self-Service Analysis Platform 104 demonstrates how the quality governance engine provides reliable data to the analysis platform, while the analysis platform uses intelligent data retrieval lists to guide the optimization of governance standards, forming an intelligent closed loop of "governance-application".
[0052] like Figure 2 The diagram shows the intelligent data collection triggering process, clearly illustrating how business events (node status updates) are automatically transformed into data actions (collection tasks). The core of steps S202-S203 lies in the automatic "listen-generate" mechanism, resolving issues of human error and delays. The feedback in step S206 forms a closed loop, ensuring that the business state and data state remain synchronized.
[0053] like Figure 3 The diagram shown is a dynamic quality audit flowchart. This flowchart highlights the depth of quality inspection. Steps S303-S304 are the core, demonstrating how to leverage knowledge graphs to leap from "static rule checking" to "dynamic semantic auditing." It can uncover complex contradictions such as "the daily fluid production of a certain oil well far exceeds the theoretical upper limit predicted by the geological model of its reservoir," which is impossible with traditional methods. Step S306 ensures that all issues are systematically tracked and managed.
[0054] The advantages and positive effects of this invention are: 1. This invention provides source protection and achieves both quality and efficiency. By using digital twin-driven data collection, it ensures the integrity and timeliness of data from the source.
[0055] 2. This invention features an intelligent deep audit function, which achieves semantic-level quality checks through a dynamic knowledge graph, significantly improving the accuracy of data and its relevance to business needs.
[0056] 3. This invention is highly automated, from data collection and triggering to data synchronization and quality inspection, greatly reducing manual intervention, lowering operation and maintenance costs and error rates.
[0057] 4. This invention features agile business empowerment, a low-code / no-code platform, and an intelligent data retrieval list, which greatly reduces the threshold for data use, quickly responds to personalized analysis needs, and directly empowers intelligent research and scientific decision-making.
[0058] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A data lake governance system based on digital twins and dynamic knowledge graphs, characterized in that: include, The business digital twin module is used to build and maintain digital twin models of business projects, and the digital twin models synchronize the status of key nodes of the business projects in real time. The intelligent acquisition triggering and verification module is used to automatically trigger and verify acquisition tasks based on the model status; The data access and dynamic quality governance engine is used to receive data from source collection tools and use a pre-built dynamic knowledge graph to perform semantic-level correlation and compliance audits on the received data. A low-code / no-code self-service analytics platform is used to respond to users' analytical operation intentions and automatically generate and recommend a list of data to be retrieved based on the dynamic knowledge graph to complete the analysis.
2. The data lake governance system based on digital twins and dynamic knowledge graphs according to claim 1, characterized in that: The data access and dynamic quality governance engine includes, The adaptive synchronizer is configured to receive data streams actively pushed by the source tools via a standardized API; The dynamic knowledge graph construction unit is used to automatically build and update domain knowledge graphs describing data entities, business rules and their relationships from enterprise data standard documents, business system metadata and historical data question databases; A semantic-level quality auditor is used to perform semantic-level correlation and compliance checks on the data entering the lake using the domain knowledge graph.
3. The data lake governance system based on digital twins and dynamic knowledge graphs according to claim 1 or 2, characterized in that: The low-code / no-code self-service analytics platform includes a visual orchestration environment, an integrated coding environment, and an intelligent data retrieval list generator. The intelligent data retrieval list generator is used to query the dynamic knowledge graph based on the user's analytical intent on the platform, automatically reason, and generate a standardized data retrieval list.
4. A data lake governance method based on digital twins and dynamic knowledge graphs, characterized by: Using the intelligent management and optimization system based on exploration big data as described in any one of claims 1 to 3, Includes the following steps, S1: Construct a digital twin model of the business project, wherein the digital twin model synchronizes the status of key nodes of the business project in real time; S2: Monitor the node status changes of the digital twin model, and when the status meets the preset conditions, automatically generate and send a data acquisition task to the designated terminal; S3: Receive data from source collection tools and use a pre-built dynamic knowledge graph to perform semantic-level correlation and compliance audits on the received data; S4: In the self-service analysis platform, in response to the user's analysis operation intention, the platform automatically generates and recommends a list of data to be retrieved to complete the analysis based on the dynamic knowledge graph.
5. The data lake governance method based on digital twins and dynamic knowledge graphs according to claim 4, characterized in that: In S2, a dynamic verification rule is built into the data acquisition task to perform preliminary format and logic verification when the data is submitted, and the task completion status is fed back to the digital twin model.
6. The data lake governance method based on digital twins and dynamic knowledge graphs according to claim 4 or 5, characterized in that: In S3, information is extracted from enterprise data standard documents, business system metadata, and historical data problem databases. Through entity recognition and relationship extraction technologies, a graph structure describing data assets and their business rules is automatically constructed and continuously updated to build the dynamic knowledge graph.
7. The data lake governance method based on digital twins and dynamic knowledge graphs according to claim 4 or 5, characterized in that: In step S3, based on the dynamic knowledge graph, entities and business rules associated with the audited data are queried, and cross-field and cross-table logical consistency checks and business threshold judgments are performed to conduct semantic-level correlation and compliance audits.
8. The data lake governance method based on digital twins and dynamic knowledge graphs according to claim 4 or 5, characterized in that: In step S4, the data retrieval list includes: the standard field names, formats, business meanings, and suggested sources for the required data.
9. A computer device comprising a memory, a processor, and an algorithm stored in the memory and executable on the processor, characterized in that: When the processor executes a computer program, it implements the data processing method as described in any one of claims 4 to 8.
10. A computer-readable storage medium storing a computer algorithm, characterized in that, When the computer algorithm is executed by the processor, it performs the data processing as described in any one of claims 4 to 8.