Enterprise digitalization-oriented data management and application method

By establishing a cross-departmental data management organization system and multi-level data standards, the shortcomings of existing data governance methods in terms of comprehensiveness and depth have been addressed. This has enabled the standardized governance and business value transformation of multi-source heterogeneous data, and improved the depth and adaptability of data processing.

CN120873040AActive Publication Date: 2025-10-31BEIJING HKRSOFT TECH CO LTD

Patent Information

Application Number
CN202510990763.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing data governance methods are insufficient in terms of comprehensiveness and depth, making it difficult to conduct cross-dimensional correlation analysis and potential value mining of massive and complex data from multiple sources and heterogeneity. This results in insufficient adaptability of data processing to business needs and an inability to effectively support load forecasting and equipment maintenance strategy optimization during the power operation phase.

Method used

Establish a cross-departmental data management organizational system covering the decision-making, organization and coordination, data management, and work execution levels. Through metadata collection, asset catalog construction, multi-level data standard system, and quality control rules, achieve standardized governance and business value transformation of multi-source heterogeneous data.

Benefits of technology

It has achieved the integrity and timeliness of multi-source heterogeneous data, improved the mining of potential correlation patterns across systems and data types, solved the problem of mismatch between data processing and business needs, and promoted the value transformation of data from collection to business application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873040A_ABST
    Figure CN120873040A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information and data processing, in particular to an enterprise digitalization-oriented data management and application method, which comprises the following steps of: acquiring multi-source heterogeneous data and recording metadata; establishing a cross-department management system to standardize the treatment process; building an asset catalog based on the business section and the data field classification; formulating unified data description of a multi-level data standard system; designing a quality control rule to filter low-quality data; constructing a business index system and embedding the business index system into a business system; identifying a cross-system and cross-type association mode through a metadata association technology and a rule engine; and dynamically adjusting the treatment process by adopting a multi-objective optimization algorithm in combination with service feedback and treatment parameters. According to the method, the defects of existing data governance in comprehensiveness, depth and business adaptability are overcome, standardized governance and potential value mining of multi-source data are achieved, and effective data support is provided for intelligent decision making of enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information and data processing technology, and in particular to a data governance and application method for enterprise digitalization. Background Technology

[0002] In the process of enterprise digitalization, data governance, as a core technical aspect of information resource management, focuses on the standardized control of the entire data lifecycle. Information processing systems encompass the entire process of data collection, storage, processing, analysis, and application. Through data processing, these systems ultimately form a traceable, associative, and reusable data asset system, supporting intelligent decision-making for enterprises.

[0003] The existing technical pain points of data governance and application for enterprise digitalization are as follows: Existing data governance methods are deficient in comprehensiveness and depth, making it difficult to conduct cross-dimensional correlation analysis and potential value mining of massive, complex, multi-source, heterogeneous data, resulting in insufficient adaptability of data processing to business needs. Taking power construction companies as an example, their business data typically includes equipment operation logs and engineering monitoring sensor data during the construction phase, as well as core data such as user electricity consumption records and grid load data during the power operation phase. Existing data governance methods can only achieve basic storage and simple cleaning through data warehouses or data lakes, failing to conduct in-depth correlation analysis of cross-system and cross-type data, such as the correlation between equipment failure frequency and peak user electricity consumption periods. Consequently, they cannot provide effective data support for business scenarios such as load forecasting and equipment maintenance strategy optimization during the power operation phase, limiting the value transformation of data in business decision-making. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a data governance and application method for enterprise digitalization, which solves the problems of insufficient comprehensiveness and depth of existing data governance methods, making it difficult to conduct cross-dimensional correlation analysis and potential value mining of massive and complex data, resulting in insufficient adaptability of data processing to business needs.

[0005] To solve the above-mentioned technical problems, the specific technical solution of the present invention is as follows: The data governance and application method for enterprise digitalization provided by this invention includes: Establish a cross-departmental data management organizational system and data management regulations covering the decision-making level, organization and coordination level, data management level, and work execution level. The data management level is responsible for formulating governance rules and supervising the compliance of data collection. Acquire heterogeneous data from multiple sources and record metadata information through the metadata acquisition module. Store the metadata information in the data warehouse as input for subsequent governance. Based on the multi-source data characteristics and business needs of the data warehouse, data assets are classified according to business segments and data domains. An asset catalog including metadata and technical parameters is constructed through metadata management tools, resulting in a three-level mapping relationship of the asset catalog. Based on the three-level mapping relationship of the asset catalog, a multi-level data standard system was established, including reference data, metadata standards, technical standards, business standards, management standards, and dictionary tables. Based on the asset catalog and data standard system, set data quality control rules at the field level and business logic level, and use the quality rule engine to perform automated quality checks and collaborative rectification on data warehouse data, outputting data that meets quality requirements. Using the data that has passed the quality inspection, a business indicator system is constructed. After version control and permission configuration are implemented through the indicator management platform, it is embedded into the functional modules of the business system or visualized and output to the decision-making terminal. Based on the asset catalog, target data is located, data is cleaned according to the data standard system, low-quality data is filtered through quality rules, metadata association technology and rule engine are used to match parameters such as timestamps, cross-system and cross-type association patterns are identified, and the association results are input into the business indicator system and output to the business scenario.

[0006] Furthermore, the data governance and application method for enterprise digitalization described in this invention establishes a cross-departmental data management organizational system covering the decision-making level, organizational coordination level, data management level, and work execution level, including: Decision-makers determine data governance goals and resource allocation based on the company's digitalization needs; The organizational coordination level coordinates the collaboration needs of business departments and IT departments based on the goals of the decision-making level. The data management layer formulates governance rules based on collaboration needs and oversees the compliance of the data collection process; The work execution layer performs data collection, data maintenance, and data feedback according to the governance rules; The responsibility boundaries of each level are defined by the role and permission matrix in the data management system.

[0007] Furthermore, the data governance and application method for enterprise digitalization described in this invention acquires multi-source heterogeneous data including: Structured data is extracted from the business system database using ETL tools, either through full extraction or incremental extraction. It connects to the API interfaces of IoT sensors and third-party platforms to collect semi-structured data in real time at a preset frequency; Collect unstructured data such as paper documents and scanned copies through file upload interfaces or OCR recognition tools, and convert them into processable formats such as text and PDF; During the acquisition of structured, semi-structured, and unstructured data, the source identifier, acquisition timestamp, data format, and data volume of the corresponding data are recorded simultaneously through the metadata acquisition module. The metadata information is then stored in the data warehouse in association with the original data.

[0008] Furthermore, the data governance and application method for enterprise digitalization described in this invention, based on the multi-source data characteristics and business needs of the data warehouse, classifies data assets according to business segments and data domains, including: The management data domain covers data from common enterprise management scenarios; The business data domain covers data from vertical business scenarios; For each type of data asset, an asset catalog is built using a metadata management tool, where metadata is inherited from the data warehouse's collection records, and technical parameters are extracted from the data warehouse data using a data exploration tool. The asset catalog is dynamically maintained as data is updated, resulting in a three-level mapping relationship for the asset catalog.

[0009] Furthermore, the data governance and application method for enterprise digitalization described in this invention, based on the three-level mapping relationship of the asset catalog, establishes a multi-level data standard system including reference data, metadata standards, technical standards, and dictionary tables, comprising: The reference data is based on the enterprise business terminology library, which defines the enumeration values ​​of business terms to unify the understanding of business terms across departments. The metadata standard specifies the data description rules based on the metadata information of the asset catalog and forms a mapping relationship with the asset catalog; The technical standard aims to unify the data format for multi-source data in data warehouses. Dictionary tables define cross-system data mapping rules; Standard version control is performed using standard document management tools.

[0010] Furthermore, the data governance and application method for enterprise digitalization described in this invention, based on an asset catalog and data standard system, sets field-level and business logic-level data quality control rules, including: Field-level rules are designed based on the technical parameters of the asset catalog, including rules for checking completeness, consistency, standardization, accuracy, and timeliness. Business logic-level rules are designed based on the enumeration values ​​of business terms and cross-system mapping rules of the data standard system, including association rules, rationality rules, and validity rules; The quality rule engine automates the scheduling and execution of data warehouse data, records quality issues in the quality issue ledger, receives information from data managers and data providers on how to rectify issues, and outputs data that meets quality requirements for building a business indicator system.

[0011] Furthermore, the data governance and application method for enterprise digitalization described in this invention utilizes data that has passed quality inspections to construct a business indicator system, including: Through business interviews, system log analysis, and historical report extraction, business indicators are extracted from the data that has passed quality inspection and categorized by business domain; the categorized indicators are then standardized with defined rules, data retrieval rules are unified, and relationships between indicators are established. Implement indicator version control and configure indicator permissions through the indicator management platform; After the indicators are calculated and generated based on the business system, they are output to the decision-making terminal through visualization tools.

[0012] Furthermore, the data governance and application method for enterprise digitalization described in this invention identifies cross-system and cross-type association patterns, including: Target data is located based on the constructed asset catalog; The target data is cleaned according to the established multi-level data standard system; Filter low-quality data through designed quality control rules; Utilize metadata association technology and rule engine to identify cross-system and cross-type association patterns; The associated results are input into the constructed business indicator system and output to the business scenario.

[0013] Furthermore, the data governance and application method for enterprise digitalization described in this invention also includes: Regularly collect feedback from business departments on the effectiveness of data application, asset catalog update records, quality issue logs, and indicator usage logs to build a multi-dimensional optimized input dataset; A multi-objective optimization algorithm is used, and the following optimization objectives are defined: Reduce the frequency of quality issues such as missing fields and logical contradictions; Shorten the response cycle for standard adjustments; Reduce the number of times indicator association rules fail; The algorithm executes by using governance rules, metadata standards, and quality rules as initial solutions to generate a population that includes candidate adjustment schemes. The performance of each set of solutions in historical business scenarios is evaluated by predictive simulation tools, the improvement effect of the optimization target is calculated, the optimal solution is selected based on ranking and crowding distance, and crossover and mutation are performed to generate a new generation of candidate solutions. Repeat the evaluation and optimization until the population converges, and output the Pareto optimal solution set; The cross-departmental management organization system reviews and selects solutions that meet the company's real-time priorities, adjusts data standards, updates quality rules, optimizes the indicator system, and simultaneously updates the asset catalog and the correlation analysis logic for identifying cross-system and cross-type association patterns.

[0014] Furthermore, the data governance and application method for enterprise digitalization described in this invention also includes: Data asset classification and multi-level data standard system work together to unify the definition and format of data in different business segments through the three-level mapping relationship of the asset catalog and the description rules of the standard system; Quality control rules are integrated with business indicator systems to filter out low-quality data through quality checks. The correlation analysis and dynamic optimization mechanism work together. The optimization mechanism adjusts governance parameters based on business feedback, driving the update of the correlation analysis logic that identifies cross-system and cross-type correlation patterns.

[0015] Beneficial effects of this invention; This invention addresses the shortcomings of existing methods in achieving comprehensive data input due to incomplete data coverage or non-standardized data collection processes. Through the coordinated design of asset classification, a multi-level standard system, and quality control rules, this invention establishes a standardized process for collecting all types of multi-source heterogeneous data, including structured, semi-structured, and unstructured data. The standardized process encompasses goal setting at the decision-making level, demand coordination at the organizational level, rule formulation at the data management level, and process execution at the work execution level. Furthermore, it achieves completeness and timeliness of data input by addressing the lack of comprehensiveness caused by incomplete data coverage or non-standardized data collection in existing methods. The invention also utilizes the collaborative design of asset classification, a multi-level standard system, and quality control rules. Asset classification includes management data domains and business data domains. The multi-level standard system includes unified terminology for reference data, standardized metadata descriptions, unified technical standard formats, and cross-system coding for dictionary table mappings. The quality control rules include field-level basic checks and business logic-level correlation verification. Combined with metadata correlation technology and deep analysis by the rule engine, it enables the mining of potential correlation patterns across systems and data types, thereby enhancing the depth of data governance. Through a dynamic optimization mechanism and closed-loop feedback of business needs, including multi-dimensional input data collection, NSGA-II multi-objective algorithm optimization, and cross-departmental review to adjust parameters, the governance process is dynamically adapted. The governance process data standards, quality rules, indicator system, and correlation logic solve the problem of data processing mismatch with business needs caused by static execution of existing methods. Ultimately, it realizes the full-link value transformation of data from collection to business application, providing effective data support for enterprise intelligent decision-making. Attached Figure Description

[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on the drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a data governance and application method for enterprise digitalization, as provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The technical solutions provided by various embodiments of this invention will be described in detail below with reference to the accompanying drawings. To better understand the objectives of this invention, it will be described in further detail below.

[0019] Please see Figure 1 The data governance and application method for enterprise digitalization provided by this invention includes: Step 1: Establish a cross-departmental data management organizational system covering the decision-making level, organization and coordination level, data management level, and work execution level. The data management level is responsible for formulating governance rules and supervising the compliance of data collection. Step 2: Acquire multi-source heterogeneous data and record metadata information through the metadata acquisition module. Store the metadata information in the data warehouse as input for subsequent governance. Step 3: Based on the multi-source data characteristics and business needs of the data warehouse, classify data assets according to business segments and data domains, and construct an asset catalog including metadata and technical parameters through metadata management tools to obtain a three-level mapping relationship of the asset catalog. Step 4: Based on the three-level mapping relationship of the asset catalog, formulate a multi-level data standard system including reference data, metadata standards, technical standards, and dictionary tables; Step 5: Based on the asset catalog and data standard system, set data quality control rules at the field level and business logic level, and use the quality rule engine to perform automated quality checks and collaborative rectification on the data warehouse data, outputting data that meets quality requirements; Step 6: Using the data that has passed the quality inspection, construct a business indicator system, and after implementing version control and permission configuration through the indicator management platform, embed it into the business system functional modules or output it visually to the decision-making terminal. Step 7: Locate target data based on the asset catalog, clean the data according to the data standard system, filter low-quality data through quality rules, use metadata association technology and rule engine to match parameters such as timestamps, identify cross-system and cross-type association patterns, input the association results into the business indicator system and output them to the business scenario.

[0020] The data governance and application method for enterprise digitalization provided by this invention achieves standardized governance and business value transformation of multi-source heterogeneous data through a full-process design encompassing data collection, management system construction, asset classification, standard setting, quality control, indicator construction, and correlation analysis. The specific technical solutions and logical relationships of each step are as follows: Step 1: Establish a cross-departmental data management organizational structure. To ensure the standardization of the governance process, a four-level management system is established, covering the decision-making, organizational coordination, data management, and execution levels: The decision-making level clarifies data governance goals (such as "improving the accuracy of power construction load forecasting") and resource allocation (such as budget and manpower) based on the company's digital strategy (e.g., the "3-year data-driven transformation plan"). The organizational coordination level (e.g., the data governance office) coordinates the collaboration needs of business departments (e.g., the power construction and maintenance department) and IT departments according to the decision-making level's goals (e.g., coordinating IT departments to grant sensor data access permissions to business departments). The data management level (e.g., the data governance team) formulates governance rules (e.g., data collection frequency, quality thresholds) based on collaboration needs and supervises the compliance of the data collection process (e.g., verifying authorization agreements for external sensor data). The execution level (e.g., data specialists and IT engineers in business departments) executes data collection (e.g., configuring ETL task parameters), maintenance (e.g., fixing sensor data transmission interruptions), and feedback (e.g., reporting collection anomalies) according to the governance rules. A role-based access control matrix (e.g., data providers can only write collected data, while data managers can read and write cleaned data) delineates the responsibilities of each level.

[0021] Step 2: Multi-source heterogeneous data acquisition and metadata recording. Targeting the multi-source data characteristics of enterprises in industries such as power construction and building, differentiated acquisition methods are used to obtain structured, semi-structured, and unstructured data: Structured data is extracted from business system databases (such as ERP, MES, CRM) using ETL tools (such as Kettle and Apache NiFi), supporting full extraction (such as daily midnight synchronization) or incremental extraction (such as capturing database change logs); Semi-structured data is collected in real-time at preset frequencies (such as second-level or hourly levels) through the API interfaces of IoT sensors (such as power construction monitoring sensors) and third-party platforms (such as meteorological data services) (e.g., JSON-formatted equipment operating parameters, XML-formatted industry policy documents); Unstructured data is collected from paper documents and scanned copies through the file upload interface of the enterprise document management system or OCR recognition tools (such as Tesseract), and converted into processable formats such as text and PDF. During the data collection process, the source system / device ID, collection timestamp (accurate to milliseconds), data format (such as CSV, JSON, PDF), and data volume (such as number of records and file size) are recorded synchronously through the metadata collection module. The metadata information is associated with the original data and stored in the data warehouse (such as HDFS or object storage) to provide the original input basis for subsequent governance processes.

[0022] Step 3, Data Asset Classification and Catalog Construction. Based on the multi-source data characteristics and business needs of the data warehouse, data assets are classified according to business segments (such as power construction and maintenance, construction projects) and data domains: the management data domain covers general management scenario data such as human resources (employee information, attendance records) and finance (reimbursement documents, contract amounts); the business data domain covers vertical business scenario data such as investment projects (project establishment documents, budget data) and engineering projects (construction logs, material consumption). For each type of data asset, an asset catalog is constructed: the metadata directly inherits the collection records from Step 2 (such as the source device ID of "Power Construction Monitoring Sensor Data" is "G001", and the storage location is "HDFS: / sensor / data"); the technical parameters are extracted from the data warehouse data through data exploration tools (such as the timestamp field of "Equipment Operation Log" is of type VARCHAR(19)). The asset catalog is dynamically maintained as the data is updated (such as the addition of power construction monitoring system data), providing a foundation for subsequent data positioning and analysis.

[0023] Step 4, Development of a multi-level data standard system. Based on the three-level mapping relationship of the asset catalog in Step 3, a multi-level standard system including reference data, metadata standards, technical standards and dictionary tables is developed: Reference data defines business term enumeration values ​​based on the enterprise business terminology library (such as the enumeration values ​​of "equipment status": running (01), downtime (02), fault (03)) to unify the cross-departmental understanding of business terms; Metadata standards standardize data description rules (such as "equipment operation log" must include timestamp (required), equipment ID (required), voltage value (numerical)) and form a mapping with the metadata information of the asset catalog; Technical standards unify data formats (such as the date field adopting "YYYY-MM-DD HH:MM:SS", and the ID number field length is fixed at 18 digits) to solve the problem of chaotic format in the data collected in Step 2; Dictionary tables define cross-system data mapping rules (such as the correspondence between the "contract status" code (01: pending signing, 02: signed) of the financial system and the "contract status" code (A: pending signing, B: signed) of the project management system). Standard version control can be achieved through standard document management tools (such as Confluence). For example, V1.0 only includes power construction industry standards, while V2.0 extends to the construction industry.

[0024] Step 5, Design and Implementation of Data Quality Control Rules. Based on the asset catalog from Step 3 and the standard system from Step 4, design field-level and business logic-level quality control rules: Field-level rules include checks on completeness (e.g., the "Equipment ID" field is mandatory), consistency (e.g., the deviation of the "Contract Amount" field value across systems is ≤0.5%), standardization (e.g., the "ID Number" field is validated to be in 18-digit format), accuracy (e.g., the "Voltage Value" field ranges from 0-1000V), and timeliness (e.g., the collection delay of "Equipment Operation Log" is ≤5 minutes); Business logic-level rules include association rules (e.g., "When the project status is 'Under Construction,' the construction permit number is mandatory"), reasonableness rules (e.g., the deviation of "User Monthly Electricity Consumption" from the historical 3-month average is ≤30%), and validity rules (e.g., "Equipment Maintenance Records" must be associated with the corresponding fault work order ID). The data warehouse data is automatically scheduled and executed through preset rules (such as triggering a full quality inspection every morning at midnight), and quality problems (such as missing fields or logical contradictions) are recorded in the quality problem log. Data managers collaborate with data providers (such as sensor equipment managers) to complete the problem rectification (such as repairing the sensor time synchronization module) and output data that meets the quality requirements for building a business indicator system.

[0025] Step 6: Construction and Output of Business Indicator System. Utilizing the data that passed the quality inspection in Step 5, through business interviews (such as communicating with the power construction and operation and maintenance department about "the impact of equipment failure"), system log analysis (such as frequently accessed "equipment failure rate" in BI tool query records), and extraction of historical reports (such as "user peak electricity consumption periods" in the annual operation and maintenance summary), extract business indicators (such as "equipment failure rate" and "user peak electricity consumption period percentage") and classify them by business domain (production, sales, operation and maintenance); standardize the definition of the classified indicators (such as "equipment failure rate = number of failures / running time"), unify the data retrieval rules (such as "running time" being taken from the difference in the timestamp field of "equipment operation log" in the asset catalog of Step 3), and establish the correlation between indicators (such as the time dimension correlation between "equipment failure rate" and "user peak electricity consumption period percentage"). This invention enables version control of indicators (e.g., labeling a new "equipment health" indicator as V2.0) and configures indicator permissions (e.g., the operations and maintenance department can only view "equipment failure rate," while management can view all indicators); it embeds indicators into business system functional modules (e.g., integrating "load forecast indicator dashboard" into the power construction and dispatch system) or outputs them to decision-making terminals (e.g., management's mobile APP), supporting real-time monitoring and analysis.

[0026] Step 7: Cross-dimensional correlation analysis and business application. Based on the asset catalog from Step 3, locate the target data (such as equipment operation logs and user electricity consumption records). Perform data cleaning according to the standard system from Step 4 (e.g., unify the time format to "YYYY-MM-DD HH:MM:SS" and use a mapping dictionary to resolve cross-system coding differences). Filter low-quality data using the quality rules from Step 5 (e.g., remove equipment logs with delays exceeding 5 minutes). Utilize metadata correlation technology (e.g., correlate equipment logs with user electricity consumption records based on equipment ID) and a rule engine (e.g., match timestamps for "equipment failure periods" and "user peak electricity consumption periods") to identify correlation patterns across systems (equipment management system and user service system) and across types (time-series data and behavioral data) (e.g., equipment failures are concentrated in the two hours before peak user electricity consumption). Input the correlation results into the business indicator system constructed in Step 6 (e.g., update the "equipment failure-peak electricity consumption correlation" indicator) and output them to business scenarios (e.g., power construction load forecasting models and equipment maintenance planning tools) to realize the transformation of data value into business decisions.

[0027] Specifically, the data governance and application method for enterprise digitalization described in this invention acquires multi-source heterogeneous data including: Structured data is extracted from the business system database using ETL tools, either through full extraction or incremental extraction. It connects to the API interfaces of IoT (sensors) and third-party platforms to collect semi-structured data in real time at a preset frequency; Collect unstructured data such as paper documents and scanned copies through file upload interfaces or OCR recognition tools, and convert them into processable formats such as text and PDF; During the acquisition of structured, semi-structured, and unstructured data, the source identifier, acquisition timestamp, data format, and data volume of the corresponding data are recorded simultaneously through the metadata acquisition module. The metadata information is then stored in the data warehouse in association with the original data.

[0028] In the data governance and application method for enterprise digitalization described in this invention, the collection of multi-source heterogeneous data achieves comprehensive acquisition of structured, semi-structured, and unstructured data through differentiated technical means, and simultaneously records metadata information, providing the original input basis for subsequent governance processes. The specific technical solution and logical relationships are as follows: Structured data collection is achieved through ETL tools. For structured data in enterprise business system databases (such as ERP, MES, and CRM systems), either full extraction or incremental extraction methods are used. Full extraction is suitable for scenarios with low data update frequency, such as when system workload is low in the early morning. ETL tools (such as Kettle and Apache NiFi) are used to perform a full synchronization of the business system database to obtain complete historical data. Incremental extraction, on the other hand, is for real-time updated business data. By capturing database change logs (such as Oracle's LogMiner and MySQL's Binlog), newly added or modified records are extracted in real time, ensuring the timeliness of data collection. For example, if employee information in an enterprise HR system is only updated a small amount daily, incremental extraction can be used; while contract data in a project management system, which requires complete daily synchronization, is extracted using full extraction.

[0029] The acquisition of semi-structured data is achieved through API interfaces connecting to IoT sensors and third-party platforms. For time-series data generated by IoT devices (such as power construction monitoring sensors and building equipment monitoring sensors), real-time acquisition is performed at a rate of seconds via API interfaces to obtain device operating parameters (such as voltage and current values) in JSON format. For standardized data provided by third-party platforms (such as meteorological data services and industry regulatory platforms), API interfaces are called at a rate of hours to collect industry policy documents or weather forecast data in XML format. For example, power construction companies can obtain real-time operating status data of power construction equipment through sensor API interfaces and periodically obtain regional meteorological data through meteorological platform API interfaces, providing multi-dimensional input for subsequent load forecasting.

[0030] Unstructured data is collected through file upload interfaces or OCR recognition tools. For unstructured data generated within the enterprise (such as project meeting minutes and user complaint texts), it is directly collected through the file upload interface of the enterprise document management system and stored as text or PDF format. For physical carrier data such as paper documents and scanned copies, optical character recognition (OCR) tools (such as Tesseract and ABBYY FineReader) are used to convert the text in the images into editable text format. For example, archived paper contract documents can be converted into PDF text using OCR tools, and scanned copies of handwritten construction logs from project sites can be recognized as structured text using OCR, solving the problem of unstructured data being difficult to process directly.

[0031] During the three types of data acquisition processes, metadata information is recorded synchronously through the metadata acquisition module. Metadata includes data source identifiers (such as the system ID of the business system database, the device ID of the IoT sensor), acquisition timestamps (accurate to milliseconds, such as "2025-05-27 08:30:15.123"), data formats (such as CSV for structured data, JSON for semi-structured data, and PDF for unstructured data), and data volume (such as the number of records for structured data and the file size for unstructured data). Metadata information is stored in a data warehouse (such as the HDFS distributed file system or cloud object storage) in a one-to-one association with the original data, enabling subsequent governance processes (such as data asset classification and quality control) to quickly locate and retrieve the original data based on the metadata. For example, the metadata of power construction monitoring sensor data records "Source Device ID=G001", "Acquisition Timestamp=2025-05-27 09:00:00", "Data Format=JSON", and "Data Volume=500 records", providing data traceability for subsequent analysis of the equipment's operating status during specific time periods.

[0032] The above-mentioned data collection methods cover all types of enterprise multi-source data (structured, semi-structured, and unstructured), and achieve the integrity and timeliness of data acquisition through differentiated technical means; the synchronous recording and associated storage of metadata provides key information support for subsequent data governance asset classification, standard setting and quality control, forming a complete logical chain from data collection to governance input.

[0033] Specifically, the data governance and application method for enterprise digitalization described in this invention establishes a cross-departmental data management organizational system covering the decision-making level, organizational coordination level, data management level, and work execution level, including: Decision-makers determine data governance goals and resource allocation based on the company's digitalization needs; The organizational coordination level coordinates the collaboration needs of business departments and IT departments based on the goals of the decision-making level. The data management layer formulates governance rules based on collaboration needs and oversees the compliance of the data collection process; The work execution layer performs data collection, data maintenance, and data feedback according to the governance rules; The responsibility boundaries of each level are defined by the role and permission matrix in the data management system.

[0034] In the data governance and application method for enterprise digitalization described in this invention, the establishment of a cross-departmental data management organizational system ensures the standardization and traceability of the data governance process through a four-level hierarchical division of responsibilities and collaboration mechanism design. The specific technical solutions and logical relationships of each level are as follows: As the highest guiding level of the governance system, the decision-making level clarifies data governance goals and resource allocation based on the company's digital strategy needs (such as a "3-year data-driven transformation plan"). For example, the decision-making level of a power construction company can set a data governance goal of "improving the accuracy of power construction load forecasting by 20%" by combining business pain points (such as insufficient accuracy of power construction load forecasting). In terms of resource allocation, the decision-making level needs to determine a special budget (such as an annual data governance budget of 5 million yuan) and human resource allocation (such as forming a data governance team of 20 people) to provide direction and resource guarantees for subsequent governance processes.

[0035] As the hub for cross-departmental collaboration, the organizational coordination layer coordinates the collaborative needs of business and IT departments based on the goals set by the decision-making layer. For example, the power construction and maintenance department needs to obtain IoT sensor data for equipment fault analysis, while the IT department needs to ensure the security of data access. The organizational coordination layer (such as the data governance office) needs to coordinate the needs of both parties and formulate a sensor data access permission scheme (such as granting read-only permissions to the maintenance department and setting access frequency limits) to achieve a balance between business needs and technical implementation.

[0036] As the entity responsible for formulating and supervising governance rules, the data management layer formulates specific governance rules based on the collaborative needs clearly defined by the organizational coordination layer and supervises their compliance. The rules include data collection frequency (e.g., data from power construction monitoring sensors must be collected at the second level), quality thresholds (e.g., the field missing rate in equipment operation logs must be ≤1%), and data storage standards (e.g., unstructured data must be uniformly stored in PDF format). Regarding compliance supervision, the data management layer needs to verify the authorization agreements for external data collection (e.g., usage rights for third-party meteorological data services) and check the configuration parameters of ETL tools (e.g., whether incremental extraction correctly captures database change logs), ensuring that the data collection process complies with the company's data security and compliance requirements.

[0037] The execution layer, as the specific implementer of governance rules, performs data collection, maintenance, and feedback according to the rules formulated by the data management layer. For example, business department data specialists need to configure ETL task parameters (such as setting a full extraction to be performed at 2 AM daily), IT engineers need to maintain the data transmission links of IoT sensors (such as fixing sensor data collection delays caused by network interruptions), and report collection anomalies (such as three consecutive collection failures) to the data management layer. The execution results of the execution layer (such as collection success rate and maintenance response time) directly affect the effectiveness of governance rules, and the execution process needs to be recorded through a logging system (such as the ELK log management platform) to achieve traceability of operations.

[0038] The role-based access control matrix clarifies the responsibilities of each level, ensuring clear accountability in the governance process. Data providers (such as business system administrators and sensor equipment managers) only have write permissions for raw data (e.g., writing collected data to the data warehouse) and cannot modify or delete stored data. Data managers (such as data cleaning engineers and quality analysts) have read and write permissions for cleaned data (e.g., correcting format errors and marking quality issues), but must comply with data access audit rules. Data users (such as business analysts and decision-makers) can only read the governed data (e.g., viewing the "equipment failure rate" metric through the indicator management platform) and are prohibited from directly modifying the underlying data. This detailed access control matrix avoids data security risks or blame-shifting issues caused by overlapping permissions.

[0039] The decision-making level's goals guide the collaboration direction of the organization and coordination level, the collaboration needs drive the data management level to formulate specific rules, the rules guide the actual operation of the work execution level, the permission matrix constrains the behavioral boundaries of each level, and finally forms an organizational guarantee system covering the entire data governance process.

[0040] Specifically, the data governance and application method for enterprise digitalization described in this invention, based on the multi-source data characteristics and business needs of data warehouses, classifies data assets according to business segments and data domains, including: The management data domain covers data from common enterprise management scenarios; The business data domain covers data from vertical business scenarios; For each type of data asset, an asset catalog is built using a metadata management tool, where metadata is inherited from the data warehouse's collection records, and technical parameters are extracted from the data warehouse data using a data exploration tool. The asset catalog is dynamically maintained as data is updated, resulting in a three-level mapping relationship for the asset catalog.

[0041] In the data governance and application method for enterprise digitalization described in this invention, data asset classification and asset catalog construction, by combining enterprise business scenarios and data characteristics, achieve standardized organization of multi-source data, providing a data positioning foundation for subsequent governance processes. The specific technical solutions and logical relationships are as follows: The business data domain covers data from vertical business scenarios, focusing on specialized data for specific business segments within an enterprise. Taking the power construction industry as an example, data for investment projects includes project initiation documents (project name, investment amount) and feasibility reports (expected returns, risk assessment); data for engineering projects includes construction logs (construction dates, progress descriptions) and material consumption records (material type, quantity used); data for design projects includes drawings (design version, technical parameters) and change records (reasons for changes, approval time), etc. This type of data is business-specific and directly supports the analytical needs of specific business processes (such as project progress tracking and design change impact assessment).

[0042] For each type of data asset in the management data domain and business data domain, an asset catalog is constructed using a metadata management tool. Metadata information is directly inherited from the data warehouse's collection records. For example, the metadata of power construction monitoring sensor data includes the source device ID (e.g., "G001"), collection timestamp (e.g., "2025-05-27 09:00:00"), data format (e.g., "JSON"), and storage location (e.g., "HDFS: / sensor / data"). Technical parameters are extracted from the data warehouse data using data exploration tools. For example, the technical parameters of "equipment operation log" data include field type (e.g., "timestamp" is of type VARCHAR(19)), field length (e.g., "device ID" is of type VARCHAR(10)), and field correlation (e.g., "voltage value" and "current value" have a linear correlation).

[0043] The asset catalog is dynamically maintained as data is updated, achieving real-time three-level mapping relationships. When an enterprise adds new data sources (such as deploying a power construction monitoring system) or existing data changes (such as updating sensor device IDs), metadata management tools (such as Apache Atlas) automatically collect metadata information from the records (e.g., the source device ID for adding "Power Construction Sensor Data" is "P001"), and re-extract technical parameters using data exploration tools to update the corresponding entries in the asset catalog. This ultimately forms a three-level mapping relationship of "Business Segment - Data Domain - Specific Data Item," such as "Power Construction Operation and Maintenance Business - Business Data Domain - Power Construction Monitoring Sensor Data" and "Construction Project Business - Business Data Domain - Construction Log Data," providing accurate data positioning basis for subsequent data standardization, quality control, and correlation analysis.

[0044] The division between management data domains and business data domains clarifies the business attributes of data, the extraction of metadata and technical parameters gives data descriptibility, dynamic maintenance ensures the timeliness of the catalog, and ultimately forms a traceable and associative asset system to support the efficient execution of subsequent governance processes.

[0045] Specifically, the data governance and application method for enterprise digitalization described in this invention, based on the three-level mapping relationship of the asset catalog, establishes a multi-level data standard system including reference data, metadata standards, technical standards, and dictionary tables, comprising: The reference data is based on the enterprise business terminology library, which defines the enumeration values ​​of business terms to unify the understanding of business terms across departments. The metadata standard specifies the data description rules based on the metadata information of the asset catalog and forms a mapping relationship with the asset catalog; The technical standard aims to unify the data format for multi-source data in data warehouses. Dictionary tables define cross-system data mapping rules; Standard version control is performed using standard document management tools.

[0046] The data governance and application method for enterprise digitalization described in this invention addresses the governance challenges caused by inconsistent terminology, chaotic formats, and cross-system coding differences in multi-source data through the collaborative design of reference data, metadata standards, technical standards, dictionary tables, and version control. This provides a unified standard for subsequent data quality control and correlation analysis. The specific technical solutions and logical relationships are as follows: The reference data defines enumerated values ​​for business terms based on the enterprise business terminology library, unifying the understanding of business terms across departments. The enterprise business terminology library is a standardized repository that integrates business terms from various departments (such as the power construction industry terminology library, which includes core terms such as "equipment status" and "project stage"). The reference data extracts frequently used business terms from it and defines their enumerated values. For example, the term "equipment status" in the power construction industry can be defined as: running (01), out of service (02), fault (03), under maintenance (04), to ensure that the operation and maintenance department, equipment management department, and data analysis department have a consistent understanding of the "fault" status; the term "project stage" in the construction industry can be defined as: project initiation (A), design (B), construction (C), acceptance (D), to avoid data statistical deviations caused by ambiguity in terms.

[0047] The metadata standard specifies data description rules based on the metadata information in the asset catalog and establishes a mapping relationship with the asset catalog. The metadata information in the asset catalog includes key information such as data source (e.g., the source device ID for "Power Construction Monitoring Sensor Data" is "G001"), storage location (e.g., "HDFS: / sensor / data"), and collection timestamp (e.g., "2025-05-27 09:00:00"). The metadata standard formulates data description rules based on this information. For example, the metadata standard for "Equipment Operation Log" requires the inclusion of a timestamp (required, accurate to milliseconds), device ID (required, format "G" + 4 digits), voltage value (required, numerical, range 0-1000V), and current value (optional, numerical, range 0-100A). Each rule corresponds one-to-one with the metadata information of the "Equipment Operation Log" data item in the asset catalog, ensuring the completeness and traceability of the data description.

[0048] The technical standard standard unifies the data format to address the formatting issues of multi-source data in data warehouses. Multi-source data in data warehouses often exhibits inconsistent formats (e.g., date fields may appear in various formats such as "2025 / 05 / 27" or "2025-05-27 08:30," and ID number fields may be 15 or 18 digits long). The technical standard solves these problems by defining a unified format. For example, date fields uniformly adopt the format "YYYY-MM-DD HH:MM:SS" (e.g., "2025-05-27 09:00:00"), ID number fields uniformly adopt the 18-digit numeric format (e.g., "110101199001011234"), and numeric fields uniformly adopt the DECIMAL(10,2) format (e.g., "voltage value = 220.50"). The implementation of the technical standard ensures that data from different sources follows a consistent format during storage and processing, avoiding increased data cleaning costs due to format differences.

[0049] Dictionary tables define cross-system data mapping rules to resolve inconsistencies in data encoding across systems. Different business systems within an enterprise (such as financial systems and project management systems) often use different codes for the same business concept, leading to difficulties in data association. For example, the "contract status" code in the financial system might be 01 (pending signing) or 02 (signed), while the "contract status" code in the project management system might be A (pending signing) or B (signed). The dictionary table needs to define the mapping rules between the two: 01→A, 02→B. Similarly, the "department code" in the human resources system (e.g., 001→Operations and Maintenance Department) and the "department code" in the ERP system (e.g., D01→Operations and Maintenance Department) can establish a correspondence through the dictionary table, enabling accurate matching of cross-system data during association analysis.

[0050] Standard version control is implemented through standard document management tools, ensuring the dynamic adaptability of standards. Standard document management tools (such as Confluence and GitLab) support version recording and rollback of standards. When a company's business expands or new data sources are added, version control can be used to update the standard content. For example, the initial version V1.0 only includes enumerated values ​​and format rules for terms such as "equipment status" and "project stage" in the power construction industry; when the company's business expands to the construction industry, version V2.0 can be released, adding enumerated values ​​for terms such as "construction status" and "material type" for the construction industry, and adjusting technical standards (such as adding date format rules for "construction logs"). Version control ensures that standards are synchronized with the company's business development, avoiding governance failures caused by outdated standards.

[0051] Reference data provides a terminology basis for metadata standards. Metadata standards and technical standards jointly regulate data description and format. Dictionary tables resolve cross-system coding differences, and version control ensures the adaptability of standards. Ultimately, a standardized system covering the entire data lifecycle is formed, providing a unified data language support for subsequent quality control, indicator construction, and correlation analysis.

[0052] Specifically, the data governance and application method for enterprise digitalization described in this invention, based on an asset catalog and data standard system, sets field-level and business logic-level data quality control rules, including: Field-level rules are designed based on the technical parameters of the asset catalog, including rules for checking completeness, consistency, standardization, accuracy, and timeliness. Business logic-level rules are designed based on the enumeration values ​​of business terms and cross-system mapping rules of the data standard system, including association rules, rationality rules, and validity rules; The quality rule engine automates the scheduling and execution of data warehouse data, records quality issues in the quality issue ledger, receives information from data managers and data providers on how to rectify issues, and outputs data that meets quality requirements for building a business indicator system.

[0053] In the data governance and application method for enterprise digitalization described in this invention, the data quality control rules are set through the collaborative design of field-level and business logic-level rules, combined with automated execution and rectification mechanisms, to ensure that data quality meets the needs of subsequent analysis. The specific technical solution and logical relationships are as follows: Field-level rules are designed based on the technical parameters of the asset catalog, focusing on checking the basic attributes of data fields. The technical parameters of the asset catalog include information such as field type (e.g., VARCHAR, DECIMAL), field length (e.g., VARCHAR(19)), and field relationships (e.g., the linear correlation between "voltage value" and "current value"). Field-level rules are designed based on this information to check the completeness, consistency, standardization, accuracy, and timeliness of the data. For example, for the "Equipment Operation Log" data in the power construction industry, the integrity rule requires that the "Equipment ID" field be filled in (the technical parameters show that this field is VARCHAR(10), not empty); the consistency rule requires that the deviation of the "Contract Amount" field value across systems be ≤0.5% (the technical parameters show that the amount field in the financial system and the project management system is DECIMAL(10,2)); the standardization rule requires that the "ID Number" field be validated in 18-digit format (the technical parameters show that this field is VARCHAR(18)); the accuracy rule requires that the "Voltage Value" field be in the range of 0-1000V (the technical parameters show that this field is DECIMAL(5,1), with a maximum value constraint); and the timeliness rule requires that the "Equipment Operation Log" collection delay be ≤5 minutes (the technical parameters show that the collection timestamp accuracy of the sensor data is in milliseconds).

[0054] Business logic-level rules are designed based on the enumerated values ​​of business terms and cross-system mapping rules of the data standard system, focusing on the logical relationships between data and the verification of business rationality. Reference data (such as the enumerated values ​​of "equipment status": running, down, fault) and dictionary tables (such as the mapping of "contract status" codes between the financial system and the project management system) of the data standard system provide the basis for rule design. For example, the association rule requires that "when the project status is 'under construction,' the construction permit number is mandatory" (the enumerated values ​​of "project status" in the reference data include "under construction," and the business logic requires that projects under construction must have a construction permit); the rationality rule requires that the deviation between "user's monthly electricity consumption" and the historical average of the past 3 months be ≤30% (by mapping historical data of user electricity consumption records through a dictionary table, and setting a reasonable fluctuation range based on business experience); the validity rule requires that "equipment maintenance records" be associated with the corresponding fault work order ID (through cross-system mapping rules, a unique correspondence between maintenance records and fault work orders is achieved).

[0055] The quality rule engine automates the scheduling and execution of data warehouse data, achieving high efficiency and traceability in quality checks. Quality rule engines (such as Informatica Data Quality) perform rule checks on structured, semi-structured, and unstructured data in the data warehouse according to preset scheduling strategies (e.g., triggering a full quality check at 2 AM daily and an incremental quality check every hour). During the check, the engine automatically records quality issues in a quality issue log, including the issue type (e.g., missing fields, logical contradictions), the problematic data item (e.g., missing timestamp field in "Equipment Operation Log"), the amount of problematic data (e.g., 10% of records have missing fields), and the source of the problem (e.g., abnormal data transmission from sensor device G001). Based on the log information, data managers (e.g., quality analysts) collaborate with data providers (e.g., sensor device managers) to rectify the issues: if the issue is a missing field, the device manager repairs the sensor time synchronization module; if the issue is a logical contradiction, the business department corrects the association rule between "Project Status" and "Construction Permit Number". After rectification, the engine re-executes the quality check, outputting data that meets quality requirements (e.g., complete fields, consistent logic, and standardized format), for use in the subsequent construction of the business indicator system.

[0056] Field-level rules ensure compliance of basic data attributes, business logic-level rules ensure the rationality of data business meaning, engine execution enables efficient checks, and collaborative rectification achieves closed-loop problem resolution. The final high-quality output data supports the accuracy and relevance of business indicators, meeting the needs of enterprises for data-driven decision-making.

[0057] Specifically, the data governance and application method for enterprise digitalization described in this invention utilizes data that has passed quality inspection to construct a business indicator system, including: Through business interviews, system log analysis, and historical report extraction, business indicators are extracted from the data that has passed quality inspection and categorized by business domain; the categorized indicators are then standardized with defined rules, data retrieval rules are unified, and relationships between indicators are established. Implement indicator version control and configure indicator permissions through the indicator management platform; The metrics can be embedded into the functional modules of the business system or output to the decision-making terminal through visualization tools.

[0058] In the data governance and application method for enterprise digitalization described in this invention, the construction of the business indicator system transforms high-quality data into quantifiable indicators that can drive business decisions through a full-process design encompassing indicator extraction, standard definition, version control, and application output. The specific technical solution and logical relationships are as follows: Indicator extraction and classification were achieved through business interviews, system log analysis, and historical report extraction. Key business indicators were identified from data that passed quality checks and categorized by business domain. Business interviews were conducted with core personnel in various business departments (such as heads of power construction and operation and maintenance departments and financial analysis supervisors) to collect their daily decision-making needs for indicators (such as "equipment failure rate" and "percentage of peak user electricity consumption"). System log analysis used BI tools to extract frequently accessed indicators from data query records (such as the top 3 most frequently queried indicators for the "number of failures" field in the equipment management system). Historical report extraction extracted mature indicators that had already been applied from documents such as the company's annual operation and maintenance summary and financial analysis report (such as "contract fulfillment rate" and "average material consumption"). The extracted indicators were categorized by business domain: the production domain included production-related indicators such as "equipment failure rate" and "capacity achievement rate"; the sales domain included sales-related indicators such as "percentage of peak user electricity consumption" and "customer complaint rate"; and the operation and maintenance domain included operation and maintenance-related indicators such as "repair response time" and "equipment health".

[0059] The standardization of indicator definitions and the establishment of correlations are achieved by clarifying the meaning of indicators, data retrieval rules, and logical relationships. Standardized definitions are established for categorized indicators; for example, "equipment failure rate" is defined as "the ratio of the number of equipment failures to the operating time," and "the percentage of peak electricity consumption by users" is defined as "the ratio of daily electricity consumption from 18:00 to 22:00 to the total daily electricity consumption." Unified data retrieval rules ensure consistency in indicator data sources (e.g., "operating time" is uniformly taken from the difference in timestamp fields of the "equipment operation log" in the asset catalog, and "total daily electricity consumption" is uniformly taken from the cumulative value field of user electricity consumption records). Correlationships are established between indicators; for example, "equipment failure rate" and "the percentage of peak electricity consumption by users" have a time-dimensional correlation (equipment failures mostly occur in the two hours before peak electricity consumption), and "average material consumption" and "construction progress" have a positive correlation (material consumption increases when construction progresses faster).

[0060] Metric version control and permission configuration are implemented through a metric management platform to ensure dynamic updates and secure access to metrics. The metric management platform (such as Tableau Server or Looker) supports metric version recording and retrospection: when new business needs arise (such as an enterprise expanding its power construction business), a new version of the metric can be published (e.g., "Power Construction Equipment Failure Rate" labeled V2.0), while older versions (e.g., "Traditional Power Construction Equipment Failure Rate" V1.0) are retained for historical analysis. Permission configuration is based on departmental responsibilities; for example, the operations and maintenance department only has viewing permissions for "Equipment Failure Rate" and "Maintenance Response Time," management has viewing and exporting permissions for all metrics, and the IT department has permission to modify metric metadata, preventing unauthorized access to sensitive data.

[0061] Indicators are embedded into business systems or visualized through technical integration to support real-time business decision-making. Key indicators are embedded into functional modules of business systems (e.g., a power construction dispatching system integrates a "load forecasting indicator dashboard" to display the "equipment failure-peak electricity consumption correlation" indicator in real time), enabling business personnel to directly obtain decision support while operating the business system. Indicators are output to decision-making terminals (e.g., management mobile apps, conference room screens) and presented in graphical form (e.g., line graphs showing equipment failure rate trends, heat maps showing peak electricity consumption area distribution) to visually demonstrate indicator dynamics. For example, power construction management can view the "power construction load forecasting accuracy" indicator in real time via a mobile app and adjust power construction dispatching strategies based on indicator changes; maintenance personnel can plan equipment maintenance in advance through the "equipment health" indicator dashboard in the business system.

[0062] Indicator extraction and classification clarify business needs; standardized definitions ensure the computability and consistency of indicators; version control and permission configuration guarantee the adaptability and security of indicators; and embedding indicators into business systems or providing visual output enables the value transformation of indicators. Ultimately, the business indicator system becomes a bridge connecting data governance and business decision-making, driving enterprises to transform into a data-driven management model.

[0063] Specifically, the data governance and application method for enterprise digitalization described in this invention identifies cross-system and cross-type association patterns, including: Target data is located based on the constructed asset catalog; The target data is cleaned according to the established multi-level data standard system; Filter low-quality data through designed quality control rules; Utilize metadata association technology and rule engine to identify cross-system and cross-type association patterns; The associated results are input into the constructed business indicator system and output to the business scenario.

[0064] The data governance and application method for enterprise digitalization described in this invention identifies cross-system and cross-type correlation patterns through a full-process design encompassing data location, cleaning, filtering, correlation analysis, and result application, enabling in-depth value mining of multi-source data. The specific technical solution and logical relationships are as follows: Based on the constructed asset catalog, the target data is located, and the data sources and scope to be analyzed are clarified. The asset catalog has completed the data organization through a three-level mapping relationship of "business segment - data domain - specific data item" (such as "power construction and maintenance business - business data domain - power construction monitoring sensor data" and "user service business - business data domain - user electricity consumption record data"). When locating the target data, the corresponding specific data items are extracted from the asset catalog according to the business analysis needs (such as exploring the correlation between equipment failure and user electricity consumption peak). For example, when a power construction company needs to analyze the relationship between equipment failure and user electricity consumption peak, it can quickly locate "equipment operation log" (storage location: HDFS: / device / log) and "user electricity consumption record" (storage location: HDFS: / user / consumption) through the asset catalog, obtain their metadata information (such as equipment ID, timestamp field) and technical parameters (such as the timestamp field type is VARCHAR(19)), and provide a data foundation for subsequent analysis.

[0065] The target data is cleaned according to the established multi-level data standard system to resolve issues of inconsistent data formats and cross-system coding differences. The multi-level data standard system defines technical standards (such as a unified date format of "YYYY-MM-DDHH:MM:SS") and dictionary tables (such as the coding mapping rules for "Contract Status" between the financial system and the project management system). These standards must be applied during the cleansing process. For example, the timestamp field in the "Equipment Operation Log" (originally formatted as "2025 / 05 / 27 09:00") is standardized to "2025-05-27 09:00:00" according to the technical standard; the "Regional Code" field in the "User Electricity Consumption Record" ("01" in the financial system, "A" in the user service system) is standardized to "01→A" through dictionary table mapping, ensuring consistency in data format and compatibility of coding across systems.

[0066] Low-quality data is filtered through designed quality control rules to ensure the reliability of the analyzed data. These rules include field-level (e.g., "Equipment ID" is mandatory, "Voltage Value" range 0-1000V) and business logic-level (e.g., "Equipment maintenance records must be associated with fault work order IDs") checks. These rules are applied during the filtering process to remove or correct low-quality data. For example, records in the "Equipment Operation Log" with a missing "Timestamp" field (required by field-level integrity rules) or abnormal records with a "Voltage Value" of 1500V (required by field-level accuracy rules ≤1000V) are marked as low-quality data and removed. Records in the "User Electricity Consumption Record" with logical contradictions between "Region Code" and "User Type" (required by business logic-level rationality rules for residential users to correspond to "01" region codes) are corrected and re-checked to retain only data that meets quality requirements.

[0067] Metadata association technology and rule engines are used to identify cross-system and cross-type association patterns and uncover potential relationships between data. Metadata association technology establishes associations between data items based on metadata information in the asset catalog (such as device ID and timestamp). For example, the "device ID" field can be used to associate "device operation log" with "device maintenance record", and the "timestamp" field can be used to associate "device operation log" with "user electricity consumption record". Rule engines (such as Drools and Siddhi) match data features according to preset rules (such as "the time difference between equipment failure period and user peak electricity consumption period is ≤2 hours") to identify association patterns. For example, the rule engine matched the "Equipment Operation Log" (which recorded the equipment failure time as "2025-05-27 17:30:00") and the "User Electricity Consumption Record" (which recorded the peak electricity consumption period as "2025-05-27 19:00:00-22:00:00") after cleaning and filtering. It found that the difference between the failure time and the peak start time was 1.5 hours, and identified the association pattern that "equipment failures are concentrated in the 2 hours before the peak user electricity consumption period".

[0068] The correlation results are input into the constructed business indicator system and output to business scenarios, realizing the transformation of data value into decision-making. The constructed business indicator system includes indicators such as "equipment failure rate" and "percentage of peak electricity consumption periods". Correlation results (such as "equipment failure-peak electricity consumption correlation") can be input into this system as new or updated indicators. For example, the correlation pattern of "equipment failures are concentrated in the 2 hours before peak electricity consumption" can be transformed into the "equipment failure-peak electricity consumption correlation" indicator (defined as "the percentage of failures where the difference between the failure time and the peak start time is ≤2 hours"). This information can then be updated to business system functional modules (such as the "load forecast dashboard" of the power construction and dispatch system) or visualization tools (such as the "equipment health analysis" interface of the management mobile APP) through the indicator management platform. Business personnel can adjust power construction and dispatch strategies (such as increasing the frequency of equipment inspections 2 hours before peak electricity consumption) or optimize equipment maintenance plans (such as arranging preventive maintenance of equipment with high failure rates in advance) based on this indicator, achieving data-driven and accurate decision-making.

[0069] The asset catalog provides navigation for data location, the standard system provides the basis for cleaning, the quality rules provide the criteria for filtering, the metadata association and rule engine provide the technical means for mining, and the indicator system and business scenarios provide the outlet for value transformation. Ultimately, it realizes the deep application of multi-source data from "storage" to "analysis" and then to "decision making", solving the problem of insufficient cross-dimensional correlation analysis in existing technologies.

[0070] Specifically, the data governance and application method for enterprise digitalization described in this invention also includes: Regularly collect feedback from business departments on the effectiveness of data application, asset catalog update records, quality issue logs, and indicator usage logs to build a multi-dimensional optimized input dataset; A multi-objective optimization algorithm is used, and the following optimization objectives are defined: Reduce the frequency of quality issues such as missing fields and logical contradictions; Shorten the response cycle for standard adjustments; Reduce the number of times indicator association rules fail; The algorithm executes by using governance rules, metadata standards, and quality rules as initial solutions to generate a population that includes candidate adjustment schemes. The performance of each set of solutions in historical business scenarios is evaluated by predictive simulation tools, the improvement effect of the optimization target is calculated, the optimal solution is selected based on ranking and crowding distance, and crossover and mutation are performed to generate a new generation of candidate solutions. Repeat the evaluation and optimization until the population converges, and output the Pareto optimal solution set; The cross-departmental management organization system reviews and selects solutions that meet the company's real-time priorities, adjusts data standards, updates quality rules, optimizes the indicator system, and simultaneously updates the asset catalog and the correlation analysis logic for identifying cross-system and cross-type association patterns.

[0071] In the data governance and application method for enterprise digitalization described in this invention, the dynamic optimization mechanism achieves adaptive adjustment of data governance parameters through a closed-loop design of multi-source data collection, multi-objective algorithm optimization, and solution implementation, thereby improving the adaptability of the governance process to changes in business needs. The specific technical solution and logical relationships are as follows: The construction of the multi-dimensional optimized input dataset is achieved through the regular collection of four key data categories. Feedback from business departments on the effectiveness of data application is obtained through questionnaires, system tracking, and other methods, including qualitative and quantitative evaluations such as indicator accuracy (e.g., improving the accuracy of "power construction load forecasting" from 70% to 75%) and business decision-making efficiency (e.g., reducing equipment maintenance plan formulation time from 3 days to 1 day). Asset catalog update records are automatically collected through metadata management tools (e.g., Apache Atlas), including new data items (e.g., "power construction monitoring system data"), modified metadata (e.g., adjusting the storage location of "equipment operation logs"), and deleted invalid data items (e.g., data from deactivated old sensors). Quality issue logs are recorded in real-time through a quality rule engine (e.g., Informatica Data Quality), including field missing rates (e.g., increasing the "equipment ID" field missing rate from 2% to 3%) and the number of logical contradictions (e.g., the number of records where "project status is 'under construction' but construction permit number is missing"). Indicator usage logs are collected through an indicator management platform (e.g., Tableau). Data is collected from the server, including the frequency of metric queries (e.g., the daily number of queries for "equipment failure rate" increased from 50 to 80) and the number of times association rules failed (e.g., the "equipment failure - peak electricity consumption correlation" rule failed twice due to data format mismatch). These four types of data are integrated quarterly into a multi-dimensional optimization input dataset, providing an objective basis for algorithm optimization.

[0072] The NSGA-II multi-objective optimization algorithm drives parameter adjustments through clearly defined optimization objectives. In addition to reducing the frequency of quality issues such as missing fields and logical contradictions, the algorithm defines two other objectives: shortening the standard adjustment response cycle (the time from the submission of business requirements to the completion of standard revision) and reducing the number of rule failures due to data mismatch or substandard quality. For example, power construction companies can specify their optimization objectives as: field missing rate ≤1%, standard adjustment response cycle ≤15 working days, and monthly rule failures ≤2. This multi-objective definition allows the algorithm to not only focus on data quality but also consider the efficiency of governance processes and the stability of indicator application.

[0073] The algorithm uses real-time governance parameters as the initial solution to generate candidate adjustment schemes. These real-time governance parameters include currently effective quality rule thresholds (e.g., a 5-minute delay threshold for "device operation log" collection), data standard versions (e.g., "device status" enumeration value V2.0), and indicator association rules (e.g., "time difference between equipment failure periods and peak electricity consumption periods ≤ 2 hours"). Using these parameters as the initial solution, the algorithm generates a population of 50 candidate schemes through random perturbation. Each scheme includes a combination of parameter adjustments (e.g., "delay threshold adjusted to 3 minutes + standard version upgraded to V3.0 + association rule time difference adjusted to 1.5 hours").

[0074] The evaluation of candidate solutions is validated in historical business scenarios using predictive simulation tools. These tools (such as data governance process simulators) take historical business data (e.g., equipment operation logs and user electricity consumption records from the past year) and candidate solutions as input, simulating the execution of the governance process: applying adjusted quality rules to check the data, performing data cleaning after the standard version upgrade, running adjusted indicator association rules, and outputting the improvement effects of each optimization objective (e.g., reducing the field missing rate from 3% to 1.5%, shortening the standard adjustment response cycle from 20 days to 12 days, and reducing the number of indicator association rule failures from 5 to 1). The evaluation results provide a quantitative basis for subsequent solution selection.

[0075] The selection of the optimal solution and the generation of a new generation of solutions are achieved through Pareto front ranking and crowding distance. Pareto front ranking filters out non-dominated solutions (i.e., solutions that cannot improve a certain objective without reducing other objectives), while crowding distance assesses the diversity of solutions (avoiding solutions from clustering in the same region). For example, if solution A reduces the field missing rate but prolongs the standard response time, and solution B shortens the standard response time but increases the number of indicator failures, both are Pareto optimal solutions. After selecting the optimal solution, a new generation of candidate solutions is generated through cross-validation (e.g., mixing the quality rules of solution A with the standard version of solution B) and mutation (e.g., randomly fine-tuning the time difference of association rules).

[0076] After the population converges, a Pareto optimal solution set is output, and the implementation plan is selected by the cross-departmental management organization system. When the optimization effect of five consecutive generations of candidate solutions does not show significant improvement (e.g., field missing rate fluctuation ≤ 0.1%, standard response cycle fluctuation ≤ 1 day), the algorithm determines that the population has converged and outputs a solution set including 5-8 Pareto optimal solutions (e.g., "reduce latency threshold to 3 minutes + upgrade standard version to V3.0 + add equipment aging coefficient index" or "increase field integrity threshold to 99% + shorten standard response cycle to 10 days + adjust association rule time difference to 1 hour"). The cross-departmental management organization system (e.g., the data governance office in conjunction with business departments and IT departments) selects the plan based on the company's real-time priorities (e.g., short-term focus on improving load prediction accuracy, long-term focus on reducing maintenance costs): if a rapid improvement in quality is needed, the plan with the most significant reduction in field missing rate is selected; if enhanced process flexibility is needed, the plan with a shortened standard response cycle is selected.

[0077] The implementation of the solution is achieved by adjusting the governance parameters and updating the relevant modules simultaneously. The selected solution needs to adjust the data standards (such as expanding the "equipment status" enumeration value to "under maintenance (04)"), update the quality rules (such as increasing the timeliness threshold of "power construction record" from 5 minutes to 3 minutes), optimize the indicator system (such as adding the "equipment aging coefficient" indicator), and update the asset catalog (such as adding the "power construction monitoring system data" entry) and the correlation analysis logic (such as including power construction equipment failure data and adjusting the "equipment failure-peak electricity consumption correlation" rule). After the adjustment, the governance process is rerun based on the new parameters, forming a closed loop of "data collection-governance-application-optimization", and continuously improving the adaptability of data governance to business needs.

[0078] Multi-dimensional input data provides a basis for optimization, multi-objective algorithms ensure the comprehensiveness of parameter adjustments, simulation evaluation ensures the feasibility of the solution, cross-departmental review reflects business needs orientation, and ultimately achieves dynamic optimization and continuous improvement of the data governance process.

[0079] Specifically, the data governance and application method for enterprise digitalization described in this invention ensures the integrity and timeliness of the original data input by standardizing the data collection process through a cross-departmental management organization system. Data asset classification and multi-level data standard system work together to unify the definition and format of data in different business segments through the three-level mapping relationship of the asset catalog and the description rules of the standard system; Quality control rules are integrated with business indicator systems to filter out low-quality data through quality checks. The correlation analysis and dynamic optimization mechanism work together. The optimization mechanism adjusts governance parameters based on business feedback, driving the update of the correlation analysis logic that identifies cross-system and cross-type correlation patterns.

[0080] In the data governance and application method for enterprise digitalization described in this invention, each core module forms a governance closed loop through a collaborative mechanism, enabling data to adapt to business needs throughout the entire process from collection to application. The specific technical solution and logical explanation of the collaborative relationship are as follows: Data collection methods and the synergy of cross-departmental management organizational systems The completeness and timeliness of data collection are achieved through standardized processes within a cross-departmental management organizational structure. The decision-making level, based on the company's digital strategy (such as a "3-year data-driven transformation plan"), clarifies data collection objectives (e.g., "covering over 90% of business system data") and resource allocation (e.g., deploying 50 IoT sensors). The organizational coordination level (e.g., the data governance office) coordinates the needs of business departments (e.g., the power construction and maintenance department) and IT departments, formulating collection priorities (e.g., prioritizing the collection of power construction monitoring sensor data) and resource allocation plans (e.g., allocating dedicated network bandwidth for key sensors). The data management level formulates specific collection rules (e.g., power construction monitoring sensor data must be collected in seconds, and business system databases must be fully extracted every morning), and monitors compliance through a log monitoring system (e.g., ELK) (e.g., checking whether ETL tasks are triggered on time and whether sensor data transmission delay is ≤5 minutes). The execution level (e.g., IT engineers) configures collection tools according to the rules (e.g., setting ETL task parameters and debugging sensor networks), and reports any abnormal issues (e.g., three consecutive collection failures) to the data management level in real time. Through the collaboration of the four-level system, the standardization of the data collection process is ensured: the needs of the business departments are transmitted to the execution level through the decision-making level, the technical capabilities of the IT department are matched with the business needs through the organizational coordination level, and the rules of the data management level are implemented through supervision, ultimately achieving the integrity (covering multiple data types) and timeliness (collecting at a preset frequency) of the raw data input.

[0081] The data asset classification and standard system is unified through the three-level mapping relationship of the asset catalog and the description of standard rules. The three-level mapping of "business segment - data domain - specific data item" in the asset catalog (such as "power construction and maintenance business - business data domain - power construction monitoring sensor data") clarifies the business ownership and technical attributes of the data (such as metadata including source device ID and storage location, and technical parameters including field type and length); the metadata standards (such as "equipment operation log" must include timestamp and device ID) and technical standards (such as the date format being unified as "YYYY-MM-DDHH:MM:SS") in the multi-level data standard system are formulated based on the technical attributes of the asset catalog to achieve consistency in data description (such as the "equipment ID" field in different business segments being defined as VARCHAR(10)); the dictionary table (such as the "contract status" coding mapping between the financial system and the project management system) achieves unified coding through the cross-system data item association of the asset catalog (such as "financial contract data" and "project contract data"). For example, the "equipment operation log" of power construction and maintenance business and the "construction log" of building project business clearly belong to different data domains through the asset catalog. However, by unifying the date format through technical standards and unifying the "status" code through dictionary tables, the data across business segments can be mutually recognized and correlated, thus solving the problem of confusing definitions of multi-source data.

[0082] Quality control rules and business indicator systems form a two-way support through data quality filtering and indicator application feedback. Quality control rules (such as mandatory "Equipment ID" at the field level and mandatory "Construction Permit Number during Project Construction" at the business logic level) perform automated checks on data warehouse data, filtering low-quality data with missing fields or logical contradictions (e.g., removing log records with missing "Equipment ID" and correcting records with contradictions between "Project Status" and "Construction Permit Number"), and outputting data that meets quality requirements (e.g., equipment operation logs with complete fields and consistent logic). The business indicator system is built upon this high-quality data (e.g., the "Equipment Failure Rate" indicator is calculated using "Number of Failures / Running Time," where "Running Time" is taken from the timestamp field after quality filtering), improving the accuracy (e.g., the failure rate calculation result is not affected by missing fields) and correlation (e.g., the "Equipment Failure and Peak Electricity Consumption Correlation" indicator is based on high-quality data with timestamp matching, avoiding correlation deviations caused by data logic contradictions). At the same time, problems discovered during indicator application (e.g., abnormal fluctuations in "Equipment Failure Rate") can be fed back to the quality control rules, driving rule optimization (e.g., increasing the accuracy threshold of the "Voltage Value" field).

[0083] The correlation analysis and dynamic optimization mechanism adapts to business needs through parameter adjustments and logic updates. The dynamic optimization mechanism regularly collects business feedback (such as "the improvement in load forecast accuracy"), asset catalog update records (such as newly added power construction monitoring system data), quality problem ledgers (such as field missing rate), and indicator usage logs (such as the number of times correlation rules failed) to construct a multi-dimensional optimization input dataset. The NSGA-II multi-objective optimization algorithm is used to generate Pareto optimal solution sets (such as "reducing the quality timeliness threshold to 3 minutes + expanding equipment status enumeration values ​​+ adding equipment aging coefficient indicators"). The cross-departmental management organization selects the solution that meets the current priority (such as focusing on improving load forecast accuracy in the short term). The adjusted data standards (such as expanding the "equipment status" enumeration values), quality rules (such as increasing the timeliness threshold of "user electricity consumption records") and indicator system (such as adding the "equipment aging coefficient" indicator) are synchronously updated to the correlation analysis logic (such as incorporating power construction equipment failure data and adjusting the time window of the "equipment failure-peak electricity consumption correlation" rule). For example, when business feedback indicates that "the improvement in load forecast accuracy is not significant," the optimization mechanism may adjust the time window for correlation analysis (from 2 hours to 1.5 hours) to make correlation pattern recognition more accurate (such as a closer correlation between equipment failure and peak electricity consumption), thereby improving the input quality of the prediction model and achieving dynamic adaptation of governance parameters to business needs.

[0084] By closely linking all aspects of data governance: data collection provides input for asset classification, the standard system provides descriptive basis for classification, quality control provides the foundation for indicator construction, and the optimization mechanism provides dynamic adjustment capability for correlation analysis, an adaptive governance system covering the entire data lifecycle is ultimately formed, solving the problem of insufficient adaptability between data processing and business needs in existing technologies.

[0085] This invention addresses the shortcomings of existing data governance methods in terms of comprehensiveness by covering all types of multi-source heterogeneous data and establishing standardized processes for cross-departmental management. The technical solution addresses the multi-source characteristics of enterprise data—structured (business system databases), semi-structured (IoT sensors, third-party platform APIs), and unstructured (paper documents, scanned copies)—by employing differentiated technologies such as ETL tool extraction, real-time API interface acquisition, and OCR recognition and conversion to achieve complete data collection. Simultaneously, a four-level management system—decision setting at the decision-making level, demand coordination at the organizational level, rule formulation at the data management level, and process execution at the work execution level—standardizes collection frequency, storage standards, and compliance supervision, preventing data omissions or collection biases and providing a comprehensive foundation of original inputs for subsequent governance.

[0086] This invention enhances the depth of data governance through the collaborative design of asset classification, standard systems, and correlation analysis. Asset classification based on business segments and data domains (management data domains covering general scenarios, business data domains covering vertical scenarios) clarifies data business attributes. A multi-level data standard system (using unified terminology for reference data, standardized metadata descriptions, unified technical standard formats, and cross-system coding for dictionary table mapping) resolves the problem of inconsistent data definitions. Quality control rules (field-level basic checks and business logic-level correlation verification) filter low-quality data. Metadata correlation technology (linking multi-source data based on device IDs, timestamps, etc.) and a rule engine (matching parameters such as time windows and logical conditions) are used to identify potential correlation patterns across systems (equipment management and user service systems) and types (time-series data and behavioral data) (such as the time correlation between equipment failure and peak electricity consumption), achieving a leap from data storage to in-depth value mining.

[0087] This invention improves the adaptability of data processing to business needs through a dynamic optimization mechanism and closed-loop feedback. It regularly collects business application feedback (e.g., indicator accuracy), asset catalog update records (e.g., newly added data items), quality issue logs (e.g., field missing rate), and indicator usage logs (e.g., rule failure counts) to construct a multi-dimensional input dataset. The NSGA-II multi-objective optimization algorithm is used to generate Pareto optimal solutions (e.g., adjusting quality thresholds, upgrading standard versions, and optimizing association rules). The cross-departmental management organization selects solutions based on real-time enterprise priorities, adjusting data standards, updating quality rules, optimizing the indicator system, and synchronously updating the association analysis logic (e.g., incorporating new data items and adjusting time windows), thus solving the adaptability issues caused by static execution in existing methods.

Claims

1. A data governance and application method for enterprise digitalization, characterized in that, include; Acquire heterogeneous data from multiple sources and collect metadata information, which is then stored in a data warehouse as input for subsequent governance. Based on the multi-source data characteristics and business needs of the data warehouse, data assets are classified according to business segments and data domains. An asset catalog including metadata and technical parameters is constructed through metadata management tools, resulting in a three-level mapping relationship of the asset catalog. Based on the three-level mapping relationship of the asset catalog, a multi-level data standard system was established, including reference data, metadata standards, technical standards, business standards, management standards, and dictionary tables. Based on the asset catalog and data standard system, set data quality control rules at the field level and business logic level, and use the quality rule engine to perform automated quality checks and collaborative rectification on data warehouse data, outputting data that meets quality requirements. Using the data that has passed the quality inspection, a business indicator system is constructed. After version control and permission configuration are implemented through the indicator management platform, it is embedded into the business system functions or visualized and output to the decision-making terminal. Based on the asset catalog, target data is located, data is cleaned according to the data standard system, low-quality data is filtered through quality rules, metadata association technology and rule engine are used to match timestamp parameters, cross-system and cross-type association patterns are identified, and the association results are input into the business indicator system and output to the business scenario.

2. The data governance and application method for enterprise digitalization according to claim 1, characterized in that, Establishing a cross-departmental data management organizational system covering the decision-making, coordination, data management, and execution levels includes: Decision-makers determine data governance goals and resource allocation based on the company's digitalization needs; The organizational coordination level coordinates the collaboration needs of business departments and IT departments based on the goals of the decision-making level. The data management layer formulates governance rules based on collaboration needs and oversees the compliance of the data collection process; The work execution layer performs data collection, data maintenance, and data feedback according to the governance rules; The responsibility boundaries of each level are defined by the role and permission matrix in the data management system.

3. The data governance and application method for enterprise digitalization according to claim 2, characterized in that, Acquiring multi-source heterogeneous data includes: Structured data is extracted from the business system database using ETL tools, either through full extraction or incremental extraction. It connects to the API interfaces of IoT sensors and third-party platforms to collect semi-structured data in real time at a preset frequency; Collect paper documents and scanned copies through file upload interfaces or OCR recognition tools, and convert them into text and PDF; During the collection of structured, semi-structured, and unstructured data, the source identifier, collection timestamp, data format, and data volume of the corresponding data are collected and recorded simultaneously. Metadata information is associated with the original data and stored in the data warehouse.

4. The data governance and application method for enterprise digitalization according to claim 3, characterized in that, Based on the multi-source data characteristics and business needs of data warehouses, data assets are classified according to business segments and data domains, including: The management data domain covers data from common enterprise management scenarios; The business data domain covers data from vertical business scenarios; For each type of data asset, an asset catalog is built using a metadata management tool, where metadata is inherited from the data warehouse's collection records, and technical parameters are extracted from the data warehouse data using a data exploration tool. The asset catalog is dynamically maintained as data is updated, resulting in a three-level mapping relationship for the asset catalog.

5. The data governance and application method for enterprise digitalization according to claim 4, characterized in that, Based on the three-level mapping relationship of the asset catalog, a multi-level data standard system was established, including reference data, metadata standards, technical standards, and dictionary tables. The reference data is based on the enterprise business terminology library, which defines the enumeration values ​​of business terms to unify the understanding of business terms across departments. The metadata standard specifies the data description rules based on the metadata information of the asset catalog and forms a mapping relationship with the asset catalog; The technical standard aims to unify the data format for multi-source data in data warehouses. Dictionary tables define cross-system data mapping rules; Standard version control is performed using standard document management tools.

6. The data governance and application method for enterprise digitalization according to claim 5, characterized in that, Based on the asset catalog and data standard system, the following data quality control rules are set at the field level and business logic level: Field-level rules are designed based on the technical parameters of the asset catalog, including rules for checking completeness, consistency, standardization, accuracy, and timeliness. Business logic-level rules are designed based on the enumeration values ​​of business terms and cross-system mapping rules of the data standard system, including association rules, rationality rules, and validity rules; The quality rule engine automates the scheduling and execution of data warehouse data, records quality issues in the quality issue ledger, receives information from data managers and data providers on how to rectify issues, and outputs data that meets quality requirements for building a business indicator system.

7. The data governance and application method for enterprise digitalization according to claim 6, characterized in that, Using the data from the quality inspection, a business indicator system is constructed, including: Through business interviews, system log analysis, and historical report extraction, business indicators are extracted from the data that has passed quality inspection and categorized by business domain; the categorized indicators are then standardized with defined rules, data retrieval rules are unified, and relationships between indicators are established. Implement indicator version control and configure indicator permissions through the indicator management platform; After the indicators are calculated and generated based on the business system, they are output to the decision-making terminal through visualization tools.

8. The data governance and application method for enterprise digitalization according to claim 7, characterized in that, Identifying cross-system, cross-type association patterns includes: Target data is located based on the constructed asset catalog; The target data is cleaned according to the established multi-level data standard system; Filter low-quality data through designed quality control rules; Utilize metadata association technology and rule engine to identify cross-system and cross-type association patterns; The associated results are input into the constructed business indicator system and output to the business scenario.

9. The data governance and application method for enterprise digitalization according to claim 8, characterized in that, Also includes: Regularly collect feedback from business departments on the effectiveness of data application, asset catalog update records, quality issue logs, and indicator usage logs to build a multi-dimensional optimized input dataset; A multi-objective optimization algorithm is used, and the following optimization objectives are defined: Reduce the frequency of quality issues such as missing fields and logical contradictions; Shorten the response cycle for standard adjustments; Reduce the number of times indicator association rules fail; The algorithm executes by using governance rules, metadata standards, and quality rules as initial solutions to generate a population that includes candidate adjustment schemes. The performance of each set of solutions in historical business scenarios is evaluated by predictive simulation tools, the improvement effect of the optimization target is calculated, the optimal solution is selected based on ranking and crowding distance, and crossover and mutation are performed to generate a new generation of candidate solutions. Repeat the evaluation and optimization until the population converges, and output the Pareto optimal solution set; The cross-departmental management organization system reviews and selects solutions that meet the company's real-time priorities, adjusts data standards, updates quality rules, optimizes the indicator system, and simultaneously updates the asset catalog and the correlation analysis logic for identifying cross-system and cross-type association patterns.

10. The data governance and application method for enterprise digitalization according to claim 9, characterized in that, Also includes: Data asset classification and multi-level data standard system work together to unify the definition and format of data in different business segments through the three-level mapping relationship of the asset catalog and the description rules of the standard system; Quality control rules are integrated with business indicator systems to filter out low-quality data through quality checks. The correlation analysis and dynamic optimization mechanism work together. The optimization mechanism adjusts governance parameters based on business feedback, driving the update of the correlation analysis logic that identifies cross-system and cross-type correlation patterns.

Citation Information

Patent Citations

  • Data management platform based on data full lifecycle management

    CN106203828A

  • A data governance driven data sharing exchange system and a working method thereof

    CN109344133A

  • Method for realizing data standard and data quality association processing based on metadata in big data governance

    CN110119395A

  • Data center station system

    CN112396404A

  • Data management system based on mine big data

    CN117874111A

Cited By

  • Real-time interaction-oriented index caliber unified processing method, computer equipment and computer program product

    CN121279962A

  • Self-service data preparation method for data weaving platform

    CN121455911A

  • Power supply enterprise marketing business data management method and system

    CN121681520A

  • Engineering construction general data management method and system

    CN121743287A

  • Data collaborative operation method for multiple service units

    CN122134308A