Enterprise data integration and insight method, device and equipment based on data-in-channel and multi-dimensional data governance and storage medium
By adopting enterprise data integration and insight methods based on data middle platform and multidimensional data governance, the inefficiency of traditional data warehouses in multi-source heterogeneous data integration scenarios is solved. It realizes efficient standardized data processing, dynamic catalog construction and multi-dimensional analysis, improves data management efficiency and security, and supports intelligent enterprise decision-making.
Patent Information
- Application Number
- CN202511925410.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-01-16
AI Technical Summary
Traditional data warehouses suffer from poor scalability, difficulty in handling unstructured data, low efficiency in multi-source heterogeneous data integration scenarios, lack of lightweight automated governance frameworks, heavy reliance on manual intervention for data standardization, inability to adapt to rapidly changing business needs, weak data preprocessing capabilities of business intelligence tools, limited analytical dimensions, and a lack of real-time anomaly detection and prediction capabilities.
Based on a data platform and multidimensional data governance, the system receives heterogeneous data from multiple sources, performs standardized processing, dynamically constructs a data asset catalog, and combines a multidimensional analysis engine for correlation analysis to generate visual insights. It also trains an anomaly detection model for real-time detection and employs dynamic permission management and privacy protection technologies.
It improves data retrieval efficiency and reusability, deeply mines data value, generates automatically visualized insights and energy consumption anomaly warnings and optimization suggestions, enhances the enterprise's intelligent decision support capabilities, and ensures the compliance and security of data throughout its entire lifecycle.
Smart Images

Figure CN121350018A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a method, apparatus, device, and storage medium for enterprise data integration and insight based on a data middle platform and multidimensional data governance. Background Technology
[0002] As enterprises deepen their digital transformation, the amount of data generated by various business systems is growing exponentially. Enterprises urgently need to improve operational efficiency, optimize resource allocation, and support intelligent decision-making through data integration and analysis. Especially in the fields of intelligent manufacturing and the Industrial Internet of Things (IIoT), the integrated analysis of multi-dimensional data such as energy management, production scheduling, and equipment monitoring has become a key means for enterprises to reduce costs and increase efficiency, creating an urgent need for full-domain data integration and intelligent insight technologies.
[0003] However, traditional data warehouse technologies suffer from poor scalability, struggle to handle unstructured data, and are inefficient in scenarios involving the integration of multi-source, heterogeneous data. While basic data platforms can achieve initial data integration, they lack a lightweight, automated governance framework, and data standardization heavily relies on manual intervention, making it difficult to adapt to rapidly changing business needs. General-purpose business intelligence tools have weak data preprocessing capabilities, cannot be deeply integrated with internal enterprise systems, resulting in limited analytical dimensions and a lack of real-time anomaly detection and prediction capabilities.
[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for enterprise data integration and insight based on a data platform and multidimensional data governance, aiming to solve the technical problem of scattered storage of internal and external data with inconsistent formats, which makes it difficult to integrate and share data.
[0006] To achieve the above objectives, this invention provides a method for enterprise data integration and insight based on a data platform and multidimensional data governance. The method includes the following steps: Receive heterogeneous data from multiple sources; The multi-source heterogeneous data is standardized to obtain a standardized dataset; A data asset catalog is dynamically constructed based on the standardized dataset; By combining a multi-dimensional analysis engine, correlation analysis is performed on the energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insights.
[0007] In one embodiment, the step of receiving multi-source heterogeneous data includes: Receives data from equipment sensors, enterprise resource planning (ERP) systems, and manufacturing execution systems. Distributed storage technology is used to collect and cache the sensor data of the equipment, the data of the enterprise resource planning system, and the data of the manufacturing execution system at high concurrency, so as to obtain integrated multi-source heterogeneous data.
[0008] In one embodiment, the step of standardizing the multi-source heterogeneous data to obtain a standardized dataset includes: The multi-source heterogeneous data is cleaned by a rule engine to remove duplicate and erroneous data, resulting in cleaned multi-source heterogeneous data. The cleaned multi-source heterogeneous data is semantically parsed and classified using semantic analysis technology to obtain cleaned and classified multi-source heterogeneous data. The cleaned and classified multi-source heterogeneous data are converted into a standardized dataset.
[0009] In one embodiment, the step of dynamically constructing a data asset catalog based on the standardized dataset includes: Based on the business attributes of the standardized dataset, a basic data catalog framework is generated; The basic data catalog framework is dynamically updated, and the catalog structure is adjusted according to the newly added data to obtain the data asset catalog. A corresponding identifier is set for each type of data in the data asset catalog, and the data is tagged to obtain a tagged data catalog.
[0010] In one embodiment, the step of combining a multi-dimensional analysis engine to perform correlation analysis on energy consumption, production, and equipment status data in the data asset catalog to obtain visual insights includes: Extract target data related to energy consumption, production, and equipment status from the data asset catalog; Perform correlation analysis on the target data to construct a multi-dimensional data model; Visual charts are generated based on the multi-dimensional data model, including energy consumption heat maps and production capacity bottleneck analysis charts. The visualized charts are output to the user interface as visual insights.
[0011] In one embodiment, the method further includes: An anomaly detection model was trained based on the energy consumption data in the standardized dataset. The aforementioned anomaly detection model is used to detect anomalies in real-time energy consumption data. When an abnormal energy consumption is detected, a corresponding energy consumption abnormality alarm message is generated; Based on the energy consumption anomaly alarm information and historical energy consumption data, energy-saving optimization suggestions are generated.
[0012] In one embodiment, the method further includes: Dynamically allocate data access permissions in the data asset catalog based on a role-based access control model; Based on the dynamically allocated permissions, sensitive data is processed using data anonymization techniques during data storage and transmission. Differential privacy technology is used to protect the privacy of the data after the anonymization process. Record data access logs and intercept unauthorized access.
[0013] Furthermore, to achieve the above objectives, this invention also proposes an enterprise data integration and insight device based on a data platform and multi-dimensional data governance, the device comprising: The data integration module is used to receive heterogeneous data from multiple sources; The data standardization module is used to standardize the multi-source heterogeneous data to obtain a standardized dataset. The catalog building module is used to dynamically build a data asset catalog based on the standardized dataset; The multi-dimensional analysis module is used to combine with the multi-dimensional analysis engine to perform correlation analysis on the energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insight results.
[0014] Furthermore, to achieve the above objectives, the present invention also proposes an enterprise data integration and insight device based on a data platform and multidimensional data governance. The device includes: a memory, a processor, and an enterprise data integration and insight program based on a data platform and multidimensional data governance stored on the memory and executable on the processor. The enterprise data integration and insight program based on a data platform and multidimensional data governance is configured to implement the steps of the enterprise data integration and insight method based on a data platform and multidimensional data governance as described above.
[0015] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing an enterprise data integration and insight program based on a data platform and multidimensional data governance. When the enterprise data integration and insight program based on a data platform and multidimensional data governance is executed by a processor, it implements the steps of the enterprise data integration and insight method based on a data platform and multidimensional data governance as described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the enterprise data integration and insight method based on data platform and multidimensional data governance as described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: By dynamically constructing a data asset catalog, data retrieval efficiency and reusability are significantly improved, while data management costs are reduced. The integrated application of multi-dimensional correlation analysis and machine learning technologies enables in-depth data value mining, automatically generating visual insights and energy consumption anomaly warnings and optimization suggestions, effectively enhancing the enterprise's intelligent decision support capabilities. A full lifecycle security control mechanism, through the comprehensive application of dynamic access control and privacy protection technologies, ensures the compliance and security of data at every stage of collection, storage, transmission, and use, comprehensively preventing data leakage risks. This achieves a synergistic improvement in data management efficiency, analytical intelligence, and security capabilities, providing systematic support for digital transformation. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating an embodiment of the enterprise data integration and insight method based on data platform and multidimensional data governance provided in this application. Figure 2 This is a system architecture diagram provided for Implementation Example 1 of the Enterprise Data Integration and Insight Method Based on Data Platform and Multidimensional Data Governance in this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the enterprise data integration and insight method based on data platform and multidimensional data governance in this application. Figure 4 This is a schematic diagram of the module structure of an enterprise data integration and insight device based on a data platform and multi-dimensional data governance according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware operating environment involved in the enterprise data integration and insight method based on data platform and multidimensional data governance in the embodiments of this application.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone; or an electronic device capable of performing the above functions, or an enterprise data integration and insight device based on a data platform and multi-dimensional data governance. The following description uses an enterprise data integration and insight device based on a data platform and multi-dimensional data governance as an example to illustrate this embodiment and the subsequent embodiments.
[0025] Based on this, the embodiments of this application provide an enterprise data integration and insight method based on a data platform and multi-dimensional data governance, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the enterprise data integration and insight method based on data platform and multidimensional data governance in this application.
[0026] In this embodiment, the enterprise data integration and insight method based on data platform and multi-dimensional data governance includes steps S10 to S40: Step S10: Receive multi-source heterogeneous data; It should be noted that the purpose of this step is to obtain raw data from various internal and external business systems of the enterprise, establish a unified data collection portal, and provide basic data resources for subsequent data processing and analysis.
[0027] Multi-source heterogeneous data refers to data sets originating from different business systems and possessing different data formats, structures, semantics, and management standards; equipment sensor data refers to monitoring data collected in real time by industrial IoT devices; enterprise resource planning (ERP) system data refers to structured business data covering core business processes such as enterprise finance, procurement, inventory, and human resources; manufacturing execution system (MES) data refers to production process data at the shop floor level, including production planning, process flow, and quality control.
[0028] Understandably, by actively polling or passively receiving data streams pushed by various data sources through a preset data access interface or IoT gateway, and after data verification and integrity checks, the acquired multi-source heterogeneous data is temporarily stored in a temporary data buffer to complete the initial collection and integration of data.
[0029] It should be understood that the purpose of this solution is to achieve automated integration and standardization of multi-source data: dynamically constructing a data asset catalog through a lightweight governance framework to reduce the cost of manual intervention; enhancing multi-dimensional data insight capabilities: introducing machine learning algorithms to achieve energy consumption anomaly detection, trend prediction, and energy-saving suggestion generation; and strengthening data security and compliance: establishing a full lifecycle security management system to ensure data privacy and controllable access permissions.
[0030] like Figure 2 As shown, the system architecture of this solution includes a data integration layer, a data processing layer, an analysis and application layer, and a security management layer. The data integration layer accesses internal and external multi-source data (such as device sensors, ERP, and MES systems) through APIs and IoT gateways; and uses distributed storage technology (such as Hadoop) to achieve high-concurrency data acquisition and caching.
[0031] The data processing layer uses a standardization module based on a rule engine and semantic analysis to automatically clean and classify data, generating standardized datasets in a unified format; the lightweight governance module dynamically builds a data asset catalog, supporting rapid retrieval and reuse of data according to business needs.
[0032] The analytics application layer uses a multi-dimensional analytics engine combined with data mining tools (such as Spark ML) to correlate data such as energy consumption, production, and equipment status to generate visual insights (such as energy consumption heatmaps and capacity bottleneck analysis); the machine learning module trains anomaly detection models (such as the Isolation Forest algorithm) to identify energy consumption anomalies in real time and push optimization suggestions.
[0033] The security control layer controls data access permissions based on the RBAC model through dynamic permission management, supporting fine-grained permission allocation; the privacy protection mechanism adopts data anonymization and differential privacy technology to ensure the compliant use of sensitive information.
[0034] In one feasible implementation, step S10 includes steps A11 to A14: Step A11: Receive equipment sensor data, enterprise resource planning system data, and manufacturing execution system data; It should be noted that the purpose of this step is to clearly define the specific source range of the received data, ensuring that it covers key data nodes in the company's core production and operation activities.
[0035] Enterprise Resource Planning (ERP) systems are management information systems that integrate the main business processes of an enterprise, and their data is usually stored in a relational database table structure; Manufacturing Execution System (MES) is a production information management system that connects the enterprise's planning layer and the shop floor control layer; Equipment sensor data refers to time-series monitoring data generated by industrial sensors such as temperature, pressure, flow, and energy consumption.
[0036] Understandably, by configuring the connection parameters and authentication information of each system, a communication link is established with the ERP database, MES database, and IoT platform. A data extraction program is used to read the target data table or subscribe to the device data topic on a timed or triggered basis. The raw data read is then encapsulated into a unified data message format to complete the targeted collection of multi-source data.
[0037] Step A12: Use distributed storage technology to collect and cache equipment sensor data, enterprise resource planning system data and manufacturing execution system data at high concurrency to obtain integrated multi-source heterogeneous data.
[0038] It should be noted that the purpose of this step is to cope with the pressure of large-scale, high-frequency data collection, improve the throughput and reliability of data access, and ensure that massive amounts of multi-source data can be stably and efficiently aggregated to the data platform.
[0039] Distributed storage technology refers to a storage architecture that distributes data across multiple physical nodes and achieves high-concurrency access through parallel processing; high-concurrency acquisition refers to processing data write requests from multiple data sources simultaneously; caching refers to using high-speed storage media as a temporary data buffer layer to alleviate backend storage pressure; integrated multi-source heterogeneous data refers to the raw data set that has been initially aggregated but has not yet undergone in-depth processing.
[0040] Understandably, by deploying a distributed message queue cluster as a data buffer layer, various types of received data are categorized by topic and written into the corresponding message queues. The distributed computing framework then consumes the data in the message queues in parallel. After simple format validation and deduplication, the data is distributed and written to a distributed file system or object storage system, thus achieving data collection and temporary storage in high-concurrency scenarios.
[0041] Step S20: Standardize the multi-source heterogeneous data to obtain a standardized dataset; It should be noted that the purpose of this step is to eliminate the differences in format, structure, and semantics between multi-source heterogeneous data, and to transform the raw data into a unified and standardized data format, so as to provide a high-quality and consistent data foundation for subsequent data governance and analysis.
[0042] Standardization processing refers to the process of cleaning, transforming, mapping, and integrating raw data according to unified data standards; a standardized dataset refers to a structured data set that meets predetermined data models and quality requirements after standardization processing.
[0043] It is understandable that by reading the raw, multi-source, heterogeneous data from distributed storage, data cleaning, semantic parsing, classification and labeling, format conversion and other processing operations are performed in sequence. The processed data is then organized according to a unified data model and written into a standardized data storage area to form a standardized dataset that can be directly accessed.
[0044] In one feasible implementation, step S20 includes steps A21 to A23: Step A21: Clean the multi-source heterogeneous data using a rule engine to remove duplicate and erroneous data, and obtain the cleaned multi-source heterogeneous data; It should be noted that the purpose of this step is to improve data quality, remove noise, redundancy and errors from the original data, and ensure the accuracy and reliability of subsequent analysis.
[0045] A rule engine refers to a computing component that automatically executes data processing logic based on predefined business rules; duplicate data refers to the situation where the same data appears multiple times due to repeated system pushes, network retransmissions, etc.; erroneous data refers to records that do not conform to business logic or data format requirements due to sensor failures, system anomalies, human error, etc.; cleaned multi-source heterogeneous data refers to a data set that has undergone quality verification and filtering, and whose data accuracy has been improved.
[0046] Understandably, by configuring data quality check rules, including but not limited to field non-empty checks, numerical range checks, time series continuity checks, and primary key duplication checks, the raw data is input into the rule engine one by one. The rule engine automatically executes the quality check logic, marking data that meets the rules as valid data and marking data that does not meet the rules as abnormal data and writing it into the abnormal log, and finally outputting the filtered and cleaned data.
[0047] Step A22: Perform semantic analysis and classification on the cleaned multi-source heterogeneous data using semantic analysis technology to obtain cleaned and classified multi-source heterogeneous data; It should be noted that the purpose of this step is to understand the business meaning of the data, realize automated data classification and labeling based on business semantics, and provide semantic support for subsequent data retrieval and correlation analysis.
[0048] Semantic analysis technology refers to the technical means of understanding data content and extracting its business meaning through natural language processing, pattern recognition and other methods; semantic parsing refers to the process of converting field names, enumeration values, descriptive text and other data in the data into standardized business terms; classification refers to classifying data according to business subject domain, data type, sensitivity level and other dimensions; cleaned and classified multi-source heterogeneous data refers to a data set that has undergone quality cleaning and semantic classification and has business readability.
[0049] Understandably, by building a business terminology library and data classification system, the cleaned data is input into a semantic analysis engine. The engine performs pattern matching and semantic mapping on the data field names and numerical meanings, identifies the business subject domains to which they belong (such as energy consumption, production, and equipment status), and assigns corresponding classification labels to the data, forming semantically labeled classified data.
[0050] Step A23: Convert the format of the cleaned and classified multi-source heterogeneous data to obtain a standardized dataset.
[0051] It should be noted that the purpose of this step is to convert data with different original formats into a standard data format, eliminate differences in data structure, and achieve physical uniformity of data.
[0052] Format conversion refers to converting the storage format, encoding method, field naming, data type, etc. of data into a form that conforms to the requirements of the target data model; a standardized dataset refers to a data set in which all data records conform to a unified data model, field definition, and encoding rules.
[0053] Understandably, by defining a standard data model and mapping rules, the cleaned and classified data is taken as input, its original fields and values are parsed item by item, the original field names are converted into target field names, the original data types are converted into target data types, the original encoded values are converted into target encoded values, and finally the converted data is reorganized according to the structure of the standard data model to form a standardized dataset in a unified format and written to the standard storage area.
[0054] Step S30: Dynamically construct a data asset catalog based on a standardized dataset; It should be noted that the purpose of this step is to establish a metadata management system for data assets, enabling unified cataloging, rapid discovery, and efficient sharing of data resources, and supporting on-demand retrieval and reuse.
[0055] A data asset catalog refers to a catalog system that describes metadata such as the attributes, structure, location, quality, and permissions of data resources; dynamic construction refers to automatically or semi-automatically updating the catalog content as data resources increase or decrease; retrieval refers to locating target data based on conditions such as keywords, tags, and business dimensions; and reuse refers to using existing data for new business scenarios.
[0056] Understandably, by scanning data resources in a standardized dataset, extracting technical metadata such as table structure, field information, data range, and update frequency, and combining this with business classification tags obtained through semantic analysis, metadata description information containing both technical and business attributes is generated. This metadata information is then entered into a data catalog repository to establish a mapping relationship between data resources and metadata, forming a queryable data asset catalog.
[0057] In one feasible implementation, step S30 includes steps A31 to A33: Step A31: Generate a basic data catalog framework based on the business attributes of the standardized dataset; It should be noted that the purpose of this step is to establish the top-level design of the data asset catalog, clarify the catalog's organizational structure and classification dimensions, and provide a framework for subsequent detailed cataloging.
[0058] Business attributes refer to the business meaning, business domain, and related business processes that the data carries; the basic data catalog framework refers to the skeleton structure of the data asset catalog, including primary classification, secondary classification and their hierarchical relationship; the catalog framework refers to the organizational form and classification standards of the catalog system.
[0059] Understandably, by parsing the semantic classification labels of each data table in the standardized dataset, the coverage of business subject domains is statistically analyzed, and a basic tree-structured directory framework is generated according to the preset business classification system (such as dividing into primary topics like energy consumption, production, and equipment status, and dividing each primary topic into secondary dimensions like data type and sensitivity level). The name, code, hierarchical relationship, and description information of each directory node are defined.
[0060] Step A32: Dynamically update the basic data catalog framework and adjust the catalog structure according to the newly added data to obtain the data asset catalog; It should be noted that the purpose of this step is to enable the data asset catalog to be adaptive, continuously reflecting changes in the company's data resources and maintaining the catalog's timeliness and completeness.
[0061] Dynamic updates refer to the automatic refresh and structural adjustment of the catalog content when new data resources are added or when existing data resources change; new data refers to new tables and fields generated by data sources that are first connected to the system or existing data sources; adjusting the catalog structure refers to adding new category nodes, merging duplicate nodes, or adjusting the node hierarchy in the existing catalog framework; the data asset catalog refers to a complete catalog that accurately describes the current status of data resources after dynamic updates.
[0062] Understandably, by monitoring the data source registration events of the data access interface or periodically scanning the metadata change logs of the standardized dataset, when a new data table or field is found, its business category tag is extracted, and it is checked whether the corresponding category node already exists in the existing directory framework. If it exists, the new data resource is mounted to the node; if it does not exist, a new category node is automatically created, and the reference relationship between the parent node and the child node is adjusted to complete the dynamic expansion and update of the directory structure.
[0063] Step A33: Set a corresponding identifier for each type of data in the data asset catalog and perform tagging processing to obtain a tagged data catalog.
[0064] It should be noted that the purpose of this step is to enhance the identifiability and retrieval of data assets, and to achieve accurate matching and flexible combination querying of data through corresponding identification and tagging systems.
[0065] The corresponding identifier refers to a globally unique code or ID within the scope of the data asset catalog, used to accurately locate a specific data resource; tagging refers to attaching multiple descriptive tags to the data resource, such as business tags, technical tags, security tags, etc.; the tagged data catalog refers to a data asset catalog in which the metadata description of the data resource contains rich tag information.
[0066] Understandably, this involves generating a globally unique identifier (such as a concatenation of "data source code_table name_field name" or a UUID) for each data table and data field, and writing it into the metadata description information. Simultaneously, based on the business classification results obtained from semantic analysis, each data resource is tagged with a business theme label; based on data type and format, a technology type label is added; and based on the results of sensitive field identification, a security level label is applied, ultimately forming a data asset catalog with unique identifiers and multi-dimensional labels.
[0067] Step S40: Combine the multi-dimensional analysis engine to perform correlation analysis on the energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insight results.
[0068] It should be noted that the purpose of this step is to break down data silos, achieve cross-business domain data fusion analysis, discover the correlation patterns and potential value between data, and provide a basis for business optimization decisions.
[0069] A multi-dimensional analysis engine refers to an analytical computing framework that can simultaneously process data from multiple business dimensions, perform cross-analysis, aggregate calculations, and pattern mining; correlation analysis refers to the analytical process of finding correlations, causal relationships, or temporal relationships among multiple datasets; and visual insight results refer to decision support information that intuitively presents the patterns, trends, anomalies, and other information obtained from the analysis in the form of charts, dashboards, etc.
[0070] Understandably, by querying and extracting data resources related to energy consumption, production, and equipment status from the data asset catalog, and integrating these data into a multi-dimensional dataset based on association keys such as timestamps, equipment IDs, and production line numbers, the multi-dimensional analysis engine performs analysis operations such as joint queries, aggregation statistics, correlation calculations, and anomaly detection. The analysis results are then mapped to visualization chart templates to generate visual insights that reflect the current business situation and problems.
[0071] In one feasible implementation, step S40 includes steps A41 to A44: Step A41: Extract target data related to energy consumption, production, and equipment status from the data asset catalog; It should be noted that the purpose of this step is to accurately locate and obtain the required data resources from massive data assets based on the analysis topic, so as to provide a data foundation for subsequent correlation analysis.
[0072] The target data refers to the specific dataset that this analysis task focuses on, namely energy consumption data, production data, and equipment status data; extraction refers to the operation process of reading the corresponding data table or field from the standardized dataset based on the metadata description information.
[0073] Understandably, by querying the tagged data catalog, data tables containing business tags such as "energy consumption," "production," and "equipment status" are retrieved, and the corresponding identifiers and storage location information of these data tables are obtained. Based on the identifier information, the corresponding data content is read from the standardized dataset, and the read data is initially aligned and integrated according to time or equipment dimensions to form the target dataset to be analyzed.
[0074] Step A42: Perform correlation analysis on the target data and construct a multi-dimensional data model; It should be noted that the purpose of this step is to establish relationships between data across business domains, forming a three-dimensional data view that can support multi-perspective analysis and reveal the inherent connections between data.
[0075] Association analysis refers to connecting or merging multiple data tables based on common dimensions (such as time, space, and equipment); multi-dimensional data model refers to a cube or star-shaped data model containing multiple business dimensions (such as time dimension, equipment dimension, and production batch dimension), which supports slicing and drilling down data from multiple angles; construction refers to the process of creating and populating a multi-dimensional data model through data modeling techniques.
[0076] Understandably, by identifying common related fields (such as timestamps and equipment numbers) in the target data, and using these fields as connection keys, data association algorithms (such as inner joins and left joins) are used to merge energy consumption data, production data, and equipment status data into a unified wide data table. The wide table data is then partitioned and indexed according to dimensions such as time, equipment, and production batches to construct a multi-dimensional data model that includes metrics (such as energy consumption, output, and equipment runtime) and multiple dimensional levels.
[0077] Step A43: Generate visualization charts based on the multi-dimensional data model, including energy consumption heat maps and production capacity bottleneck analysis charts; It should be noted that the purpose of this step is to transform the abstract data model into an intuitive graphical representation, enabling business users to quickly understand the business meaning behind the data and discover potential problems and opportunities.
[0078] Visual charts refer to a graphical display format that encodes data information through graphic elements (such as color, size, and position); energy consumption heatmaps refer to charts that display the energy consumption intensity of different equipment and different time periods in matrix form, with the intensity of color indicating the level of energy consumption; capacity bottleneck analysis charts refer to charts that display the capacity utilization rate of each link in the production process through Gantt charts, flowcharts, or bar charts, and identify the bottleneck links that limit the overall output.
[0079] Understandably, by extracting data from a multi-dimensional data model, including time, equipment, and energy consumption, and mapping energy consumption metrics to a color gradient with time (hours / days) on the horizontal axis and equipment on the vertical axis, an energy consumption heatmap is generated. Simultaneously, time and output data for each stage of the production process are extracted, the capacity utilization rate of each stage is calculated, and the utilization rate comparison of each stage is displayed in a flowchart format. Bottleneck stages with near-saturation utilization are highlighted, generating a capacity bottleneck analysis chart.
[0080] Step A44: Output the visualization chart as a visualization insight result to the user interface.
[0081] It should be noted that the purpose of this step is to deliver the analysis results to the end users, enabling business personnel to easily access and view the analysis results, supporting daily monitoring and decision-making.
[0082] The user interface refers to the graphical interface for business users to interact with, usually a web page or mobile application; output to the user interface refers to the process of embedding the generated charts into the front-end display components and providing interactive functions (such as filtering and drill-down).
[0083] Understandably, by encapsulating the generated visualization charts into visualization components, embedding these components into pre-built data analysis portals or dashboard pages, configuring data refresh frequency and interaction parameters (such as chart linkage and drill-down dimensions), when business users access the user interface, the chart data is retrieved from the backend service for rendering and display, and filter conditions are provided for users to customize the viewing scope.
[0084] Furthermore, this plan also includes: An anomaly detection model was trained based on energy consumption data in a standardized dataset. Anomaly detection models are used to detect anomalies in real-time energy consumption data. When an abnormal energy consumption is detected, a corresponding energy consumption abnormality alarm message is generated; Based on energy consumption anomaly alarm information and historical energy consumption data, energy-saving optimization suggestions are generated.
[0085] It should be noted that the purpose of this series of steps is to build a complete intelligent closed-loop management system from energy consumption anomaly detection to energy-saving suggestion generation. Through machine learning technology, it can automatically identify energy consumption anomalies and transform abnormal events into actionable optimization strategies, thereby improving the initiative and accuracy of enterprise energy management.
[0086] Anomaly detection models refer to machine learning models trained on historical data that can identify behaviors that deviate from normal energy consumption patterns. They are typically built using unsupervised learning algorithms (such as Isolation Forest). Real-time energy consumption data refers to the latest energy consumption monitoring data stream collected by sensors in real time and transmitted to the system immediately. Energy consumption anomalies refer to data points or time periods where there is a significant deviation between actual energy consumption and normal energy consumption predicted based on historical patterns. Alarm information refers to formatted notification messages that include elements such as abnormal equipment identification, occurrence time, and anomaly severity. Historical energy consumption data refers to long-term accumulated energy consumption monitoring records used for model training and root cause analysis. Energy-saving optimization suggestions refer to specific improvement measures proposed for specific abnormal events, such as adjusting equipment operating parameters, optimizing production scheduling, and equipment maintenance.
[0087] Understandably, the process begins by extracting historical energy consumption data containing features such as device ID, timestamp, energy consumption value, load rate, and ambient temperature and humidity from a standardized dataset. The data is then sliced into time windows and processed using feature engineering to construct a training sample set. This sample set is then input into the Isolation Forest algorithm for iterative training. The model parameters are adjusted until the preset detection accuracy and recall metrics are achieved on the validation set. Finally, the model file is saved and the service deployment is completed.
[0088] The system receives energy consumption data uploaded by devices in real time through a streaming data consumption interface. It performs the same feature extraction operation as in the training phase on the received data to form a feature vector, which is then input into the deployed anomaly detection model. The model outputs an anomaly score for each data point, and when the score exceeds a preset threshold, it is determined to be an anomaly.
[0089] When an anomaly is detected, the corresponding equipment identifier, anomaly occurrence time, energy consumption deviation range, and potentially affected process steps are immediately extracted. An alarm message is assembled according to a preset alarm template, written into a message queue, and pushed to energy management personnel through channels such as SMS, email, and system notifications.
[0090] After receiving an alarm, energy management personnel or automated decision-making systems query the recent historical energy consumption data of the equipment and the operating data of related equipment, analyze the abnormal patterns and trend characteristics, call the built-in energy-saving strategy knowledge base for rule matching, generate an energy-saving optimization suggestion report containing problem diagnosis, cause speculation, optimization measures, and expected energy-saving effects, and push the suggestions to the equipment control system or production scheduling system to guide on-site personnel to perform optimization operations, completing closed-loop management from anomaly discovery to problem resolution.
[0091] In practice, data flows from multiple source systems into the integration layer, is standardized, and then stored in the data lake; the governance module dynamically updates the data asset catalog, and business users can access data on demand through a visual interface; the analysis engine combines real-time and historical data to generate multi-dimensional reports, and the machine learning model simultaneously outputs anomaly alerts and optimization suggestions; the security module monitors the data flow throughout the process, intercepts unauthorized access, and records operation logs.
[0092] This lightweight data governance approach includes dynamic catalog construction rules, standardized processing workflows, and automated classification algorithms. The multi-dimensional analysis model implements correlation analysis methods and visualization logic for energy consumption, production, and equipment status data. The security mechanism design encompasses specific implementation plans for data anonymization strategies, permission allocation models, and privacy protection technologies.
[0093] This embodiment provides an enterprise data integration and insight method based on a data platform and multidimensional data governance. Based on a lightweight governance framework, it automatically builds an extensible data catalog that supports rapid retrieval and reuse; combined with a rule engine and semantic analysis, it achieves automated cleaning and format unification of heterogeneous data; it identifies energy consumption anomalies in real time through algorithms such as isolated forests and generates energy-saving optimization strategies; and it integrates dynamic permission management and differential privacy technology to ensure compliance throughout the entire process from data collection to destruction.
[0094] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 After step S40, steps S401 to S404 are also included: Step S401: Dynamically allocate data access permissions in the data asset catalog based on the role-based access control model; It should be noted that the purpose of this step is to achieve refined and dynamic management of data access permissions, ensuring that users can only access the authorized data range that matches their current business role, thereby eliminating the risk of unauthorized access and data abuse at the source.
[0095] Role-based access control (RBAC) is an access control framework that grants permissions to roles rather than directly to users. Users indirectly gain permissions by assuming roles. Roles typically correspond to positions or responsibilities within an organization, such as energy administrators, production line supervisors, and data analysts. A data asset catalog refers to a catalog system that records metadata about all data resources of an enterprise, including descriptive information such as data tables, fields, sensitivity levels, and storage locations. Dynamic allocation means that permissions are not statically granted once when a user logs in, but are calculated and granted in real time each time a data access request is initiated, based on factors such as the user's authentication status, the set of roles they belong to, the attributes of the requested data resource, and the current business context (such as access time and access terminal).
[0096] Understandably, by deploying a permission verification filter at the data service interface layer, this filter intercepts all data access requests sent to the data asset catalog, extracts the user identity token from the request message, parses the token to obtain the user identifier, queries the user-role mapping table to obtain the set of roles currently active for the user, calculates the operable permissions (including the range of accessible data tables, the list of visible fields, row-level data filtering conditions, etc.) for the target data resource based on the role-permission mapping rules, dynamically generates instructions containing the above permission policies and binds them to the user session context. Subsequent data processing logic is executed based on this dynamic permission instruction, completing the real-time allocation and enforcement of permissions.
[0097] Step S402: Based on dynamically allocated permissions, sensitive data is processed using data desensitization technology during data storage and transmission. It should be noted that the purpose of this step is to transform sensitive data fields that exceed the user's permission scope during the static storage and dynamic transmission of data. This significantly reduces the risk of sensitive information leakage in storage media and network links while ensuring the continuity of business analysis, achieving a balance between data availability and security. Dynamically allocated permissions refer to the permission instructions generated in real-time in step S401, which include a list of user-accessible fields and a list of invisible sensitive fields. Sensitive data refers to data fields involving corporate trade secrets, user privacy, or national security, which may cause damage if leaked, such as equipment operating parameters, cost data, and employee identity information. Data anonymization technology refers to a set of technical means to transform sensitive data fields, including but not limited to mask replacement, format-preserving encryption (keeping the data format unchanged but encrypting the content), and data generalization (replacing precise values with range values, such as generalizing the age "28" to "25-30"). Understandably, before data is persisted to the database or sent to the user terminal, the data processing engine reads the dynamic permission instructions bound in the current user session, parses out the list of sensitive fields that the user is not authorized to view in plaintext, calls the corresponding desensitization algorithm function for these fields, performs masking, encryption or generalization transformation processing on the original field values, and writes the transformed pseudo-values into storage media or network packets, ensuring that sensitive information always exists in desensitized form throughout the entire data lifecycle, while authorized users can view the plaintext through decryption or restoration mechanisms at the application layer, thus completing dynamic desensitization protection during data storage and transmission.
[0098] Step S403: Differential privacy technology is used to protect the privacy of the anonymized data; It should be noted that the purpose of this step is to further enhance the strength of privacy protection on the basis of de-identification. Through mathematical noise mechanism, it is ensured that even if an attacker makes multiple queries or gains some background knowledge, they will not be able to accurately reconstruct individual data records, thus meeting the strict data compliance audit and privacy protection requirements. It is particularly suitable for aggregated statistics scenarios that are published to the public or accessed by users with low privileges.
[0099] Differential privacy technology refers to a formal mathematical privacy protection method that adds random noise that conforms to a specific probability distribution to data analysis results or datasets, so that the impact of the existence or absence of any single data record on the overall query results is controlled within a preset privacy budget. Its core is the noise mechanism and privacy budget management. The privacy budget refers to the upper limit parameter for measuring the risk of privacy leakage. The smaller the budget value, the more noise is added, the higher the privacy protection strength, but the lower the data availability.
[0100] Understandably, when receiving an aggregate query request for anonymized data (such as "query the average energy consumption of a workshop over the past week"), the system identifies the size and sensitivity of the dataset involved in the query. Based on the current remaining privacy budget, it invokes a differential privacy algorithm (such as the Laplace mechanism) to calculate the intensity of random noise to be added. Random noise sampled from the Laplace distribution is then superimposed on the true aggregate result (such as the true average value), and the perturbed statistical result is returned to the queryer. Since each query consumes privacy budget, the system needs to accumulate and control the total consumption to not exceed a preset limit, thereby achieving quantifiable privacy enhancement protection and ensuring that even if an attacker initiates multiple queries, they cannot infer the original data value of an individual with high confidence.
[0101] Step S404: Record data access logs and intercept unauthorized access behavior.
[0102] It should be noted that the purpose of this step is to build a complete post-event audit and traceability and pre-event real-time protection capability, to ensure that all access behaviors are traceable and accountable, and to block and alert on unauthorized operations in a timely manner, forming a closed loop of full lifecycle security control from permission verification and data protection to audit interception, so as to meet the requirements of compliance audit and security operation.
[0103] Data access logs refer to audit trail records that record elements such as user identity, access timestamp, requested resource path, operation type (query / modify / delete), authorization result, and response status code; unauthorized access behavior refers to operations that violate security policies, such as users attempting to access data resources beyond their dynamically allocated permissions, using forged identity tokens, or attempting to bypass permission filters; interception refers to terminating the execution of a request and returning an error response before the request reaches the actual data resource.
[0104] Understandably, embedding logging functionality into the permission verification filter allows for the asynchronous extraction of key request and response information (such as user ID, target data table, time, and authorization result) for each request passing through the filter, regardless of authorization status. This information is then written to an immutable log storage system (such as blockchain or WORM storage). Simultaneously, during permission verification, if the filter determines that the requested resource exceeds the current user's dynamic permission scope or detects a missing valid identity token in the request, it immediately terminates further processing of the request, returns an error response to the front end, and pushes an alert message containing the attacker's IP address, request details, and a timestamp to the security management center. This completes real-time interception and auditing of unauthorized access, forming a comprehensive security loop encompassing pre-emptive protection, in-process control, and post-event auditing.
[0105] This embodiment provides an enterprise data integration and insight method based on a data middle platform and multi-dimensional data governance. By dynamically allocating permissions, de-identifying data, providing differential privacy protection and access blocking, it constructs a closed loop of full lifecycle security management and enhances data security protection and compliance operation capabilities.
[0106] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the enterprise data integration and insight method based on data middle platform and multidimensional data governance. Any simple modifications based on this technical concept are within the protection scope of this application.
[0107] This application also provides an enterprise data integration and insight device based on a data middle platform and multi-dimensional data governance. Please refer to [link / reference]. Figure 4 Enterprise data integration and insight devices based on data middle platform and multi-dimensional data governance include: Data integration module 10 is used to receive heterogeneous data from multiple sources; Data standardization module 20 is used to standardize multi-source heterogeneous data to obtain a standardized dataset; Catalog building module 30 is used to dynamically build a data asset catalog based on a standardized dataset; The multi-dimensional analysis module 40 is used to combine with the multi-dimensional analysis engine to perform correlation analysis on energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insights.
[0108] The enterprise data integration and insight device based on data platform and multidimensional data governance provided in this application, employing the enterprise data integration and insight method based on data platform and multidimensional data governance described in the above embodiments, can solve the technical problem of scattered storage and inconsistent formats of internal and external data, leading to difficulties in data integration and sharing. Compared with the prior art, the beneficial effects of the enterprise data integration and insight device based on data platform and multidimensional data governance provided in this application are the same as those of the enterprise data integration and insight method based on data platform and multidimensional data governance provided in the above embodiments. Furthermore, other technical features in the enterprise data integration and insight device based on data platform and multidimensional data governance are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0109] In one embodiment, the data integration module 10 is further configured to receive equipment sensor data, enterprise resource planning system data, and manufacturing execution system data; Distributed storage technology is used to collect and cache equipment sensor data, enterprise resource planning system data, and manufacturing execution system data with high concurrency, resulting in integrated multi-source heterogeneous data.
[0110] In one embodiment, the data standardization module 20 is further used to perform data cleaning on multi-source heterogeneous data through a rule engine, remove duplicate and erroneous data, and obtain cleaned multi-source heterogeneous data; Semantic analysis techniques are used to perform semantic parsing and classification on the cleaned multi-source heterogeneous data to obtain cleaned and classified multi-source heterogeneous data. The format of the cleaned and classified multi-source heterogeneous data is converted to obtain a standardized dataset.
[0111] In one embodiment, the directory building module 30 is further configured to generate a basic data directory framework based on the business attributes of the standardized dataset; The basic data catalog framework is dynamically updated, and the catalog structure is adjusted according to the newly added data to obtain the data asset catalog. Assign a corresponding identifier to each type of data in the data asset catalog and perform tagging processing to obtain a tagged data catalog.
[0112] In one embodiment, the multi-dimensional analysis module 40 is also used to extract target data related to energy consumption, production, and equipment status from the data asset catalog; Perform correlation analysis on the target data to construct a multi-dimensional data model; Visual charts are generated based on a multi-dimensional data model, including energy consumption heat maps and capacity bottleneck analysis charts. Output visual charts as visual insights to the user interface.
[0113] In one embodiment, the multi-dimensional analysis module 40 is also used to train an anomaly detection model based on energy consumption data in a standardized dataset; Anomaly detection models are used to detect anomalies in real-time energy consumption data. When an abnormal energy consumption is detected, a corresponding energy consumption abnormality alarm message is generated; Based on energy consumption anomaly alarm information and historical energy consumption data, energy-saving optimization suggestions are generated.
[0114] In one embodiment, the multi-dimensional analysis module 40 is also used to dynamically allocate data access permissions in the data asset catalog based on the role access control model; Based on dynamically allocated permissions, sensitive data is processed using data anonymization techniques during data storage and transmission. Differential privacy technology is used to protect the privacy of the anonymized data; Record data access logs and intercept unauthorized access.
[0115] This application provides an enterprise data integration and insight device based on a data platform and multidimensional data governance. The enterprise data integration and insight device based on a data platform and multidimensional data governance includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the enterprise data integration and insight method based on a data platform and multidimensional data governance in the first embodiment described above.
[0116] The following is for reference. Figure 5 This document illustrates a structural diagram of an enterprise data integration and insight device based on a data platform and multidimensional data governance, suitable for implementing embodiments of this application. The enterprise data integration and insight device based on a data platform and multidimensional data governance in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast acquisition devices, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The enterprise data integration and insight device based on data platform and multidimensional data governance shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0117] like Figure 5As shown, an enterprise data integration and insight device based on a data platform and multidimensional data governance may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the enterprise data integration and insight device based on a data platform and multidimensional data governance. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows enterprise data integration and insight devices based on data platforms and multidimensional data governance to exchange data wirelessly or via wired communication with other devices. Although the figure shows an enterprise data integration and insight device based on data platforms and multidimensional data governance with various systems, it should be understood that implementing or having all the systems shown is not required. More or fewer systems can be implemented alternatively.
[0118] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0119] The enterprise data integration and insight device based on data platform and multidimensional data governance provided in this application, employing the enterprise data integration and insight method based on data platform and multidimensional data governance described in the above embodiments, can solve the technical problem of scattered storage and inconsistent formats of internal and external data, leading to difficulties in data integration and sharing. Compared with the prior art, the beneficial effects of the enterprise data integration and insight device based on data platform and multidimensional data governance provided in this application are the same as those of the enterprise data integration and insight method based on data platform and multidimensional data governance provided in the above embodiments. Furthermore, other technical features of this enterprise data integration and insight device based on data platform and multidimensional data governance are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0120] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0122] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the enterprise data integration and insight method based on data platform and multi-dimensional data governance in the above embodiments.
[0123] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0124] The aforementioned computer-readable storage medium may be included in an enterprise data integration and insight device based on a data platform and multidimensional data governance; or it may exist independently and not be assembled into an enterprise data integration and insight device based on a data platform and multidimensional data governance.
[0125] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by an enterprise data integration and insight device based on a data platform and multidimensional data governance, the device enables the following: receiving multi-source heterogeneous data; standardizing the multi-source heterogeneous data to obtain a standardized dataset; dynamically constructing a data asset catalog based on the standardized dataset; and combining a multi-dimensional analysis engine to perform correlation analysis on energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insight results.
[0126] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0128] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0129] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the aforementioned enterprise data integration and insight method based on data platform and multidimensional data governance. This solves the technical problem of fragmented storage and inconsistent formats of internal and external enterprise data, making data integration and sharing difficult. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the enterprise data integration and insight method based on data platform and multidimensional data governance provided in the above embodiments, and will not be elaborated upon here.
[0130] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the enterprise data integration and insight method based on data platform and multidimensional data governance as described above.
[0131] The computer program product provided in this application can solve the technical problem of scattered storage and inconsistent formats of internal and external data within an enterprise, which makes data integration and sharing difficult. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the enterprise data integration and insight method based on data platform and multi-dimensional data governance provided in the above embodiments, and will not be repeated here.
[0132] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for enterprise data integration and insight based on data middle platform and multidimensional data governance, characterized in that, The method includes: Receive heterogeneous data from multiple sources; The multi-source heterogeneous data is standardized to obtain a standardized dataset; A data asset catalog is dynamically constructed based on the standardized dataset; By combining a multi-dimensional analysis engine, correlation analysis is performed on the energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insights.
2. The method as described in claim 1, characterized in that, The step of receiving multi-source heterogeneous data includes: Receives data from equipment sensors, enterprise resource planning (ERP) systems, and manufacturing execution systems. Distributed storage technology is used to collect and cache the sensor data of the equipment, the data of the enterprise resource planning system, and the data of the manufacturing execution system at high concurrency, so as to obtain integrated multi-source heterogeneous data.
3. The method as described in claim 1, characterized in that, The step of standardizing the multi-source heterogeneous data to obtain a standardized dataset includes: The multi-source heterogeneous data is cleaned by a rule engine to remove duplicate and erroneous data, resulting in cleaned multi-source heterogeneous data. The cleaned multi-source heterogeneous data is semantically parsed and classified using semantic analysis technology to obtain cleaned and classified multi-source heterogeneous data. The cleaned and classified multi-source heterogeneous data are converted into a standardized dataset.
4. The method as described in claim 1, characterized in that, The step of dynamically constructing a data asset catalog based on the standardized dataset includes: Based on the business attributes of the standardized dataset, a basic data catalog framework is generated; The basic data catalog framework is dynamically updated, and the catalog structure is adjusted according to the newly added data to obtain the data asset catalog. A corresponding identifier is set for each type of data in the data asset catalog, and the data is tagged to obtain a tagged data catalog.
5. The method as described in claim 1, characterized in that, The step of combining a multi-dimensional analysis engine to perform correlation analysis on energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insights includes: Extract target data related to energy consumption, production, and equipment status from the data asset catalog; Perform correlation analysis on the target data to construct a multi-dimensional data model; Visual charts are generated based on the multi-dimensional data model, including energy consumption heat maps and production capacity bottleneck analysis charts. The visualized charts are output to the user interface as visual insights.
6. The method as described in claim 1, characterized in that, The method further includes: An anomaly detection model was trained based on the energy consumption data in the standardized dataset. The aforementioned anomaly detection model is used to detect anomalies in real-time energy consumption data. When an abnormal energy consumption is detected, a corresponding energy consumption abnormality alarm message is generated; Based on the energy consumption anomaly alarm information and historical energy consumption data, energy-saving optimization suggestions are generated.
7. The method as described in claim 1, characterized in that, The method further includes: Dynamically allocate data access permissions in the data asset catalog based on a role-based access control model; Based on the dynamically allocated permissions, sensitive data is processed using data anonymization techniques during data storage and transmission. Differential privacy technology is used to protect the privacy of the data after the anonymization process. Record data access logs and intercept unauthorized access.
8. An enterprise data integration and insight device based on a data middle platform and multidimensional data governance, characterized in that, The device includes: The data integration module is used to receive heterogeneous data from multiple sources; The data standardization module is used to standardize the multi-source heterogeneous data to obtain a standardized dataset. The catalog building module is used to dynamically build a data asset catalog based on the standardized dataset; The multi-dimensional analysis module is used to combine with the multi-dimensional analysis engine to perform correlation analysis on the energy consumption, production, and equipment status data in the data asset catalog to obtain visualized insight results.
9. An enterprise data integration and insight device based on a data middle platform and multidimensional data governance, characterized in that, The device includes: a memory, a processor, and an enterprise data integration and insight program based on data platform and multidimensional data governance, stored on the memory and executable on the processor, wherein the enterprise data integration and insight program based on data platform and multidimensional data governance is configured to implement the steps of the enterprise data integration and insight method based on data platform and multidimensional data governance as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores an enterprise data integration and insight program based on a data platform and multidimensional data governance. When the processor executes the enterprise data integration and insight program based on a data platform and multidimensional data governance, it implements the steps of the enterprise data integration and insight method based on a data platform and multidimensional data governance as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Enterprise data platform for multi-source heterogeneous data
CN119808906A
Enterprise carbon emission analysis method and system based on ESG comprehensive evaluation model
CN120235484A
Enterprise operation decision-making method, system and equipment based on large model
CN121119742A