An AI Fusion Governance Method Based on Data Lake
Through AI technology, data is identified and integrated in the data lake, and structured models that are in line with enterprise business are generated, which solves the problem of screening limitations caused by data type diversity and achieves efficient data governance.
Patent Information
- Application Number
- CN202211651578.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-21
AI Technical Summary
In the prior art, the diversity of data types stored in the data lake leads to the screening data being limited by the enterprise's structured data model, which makes it difficult to efficiently adapt to the enterprise's business needs.
Image recognition, text recognition, and speech recognition are carried out through AI technology, structured and unstructured data are collected, and new structured models are generated that are in line with enterprise business through knowledge graph and graph database technology, and data quality evaluation is carried out in combination with supervised learning and deep learning.
The structured model of adaptive learning production is realized, the efficiency and quality of data screening are improved, and the business needs of enterprises are met.
Smart Images

Figure CN115809235B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data governance, and specifically relates to an AI fusion governance method based on a data lake. Background Art
[0002] In the past, data governance required professional technical and management personnel to operate, with relatively high threshold requirements for practical applications. Currently, the perfect integration of artificial intelligence and data governance has opened a new stage of intelligent data governance. Through AI empowerment, the operability of data governance tools can be continuously improved, enabling participants in data governance to use data governance tools more conveniently.
[0003] In the prior art, a data lake is a repository and processing system that can accommodate a large amount of raw data and has become an important tool for enterprises to apply big data. Since the data lake obtains raw data from multiple data sources of an enterprise, and for different purposes, the same raw data may have many data structures that meet specific internal model formats. Therefore, the data processed in the data lake may be any type of information, from structured data to unstructured data. These data are not all fully applicable to the structured data model of the enterprise, resulting in the filtered data being largely limited by the structured data model of the enterprise and the input enterprise business data. In view of the above situation, an AI fusion governance method based on a data lake is designed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides an AI fusion governance method based on a data lake, which has the advantage of an adaptively learned and generated structured model.
[0005] To achieve the above object, the present invention provides the following technical solution: An AI fusion governance method based on a data lake, and the steps of this method are as follows:
[0006] S1: Connect the data lake data. Automatically perform image recognition, text recognition, and speech recognition on the data connected to the data lake through AI technology, so as to collect various structured data and unstructured data through AI;
[0007] S2: Compare the collected structured data and unstructured data with the structured data model of the enterprise, and filter out the structured data that conforms to the enterprise;
[0008] S3: Integrate the data that does not conform to the structured data of the enterprise, collect metadata within the integrated data, and then perform data supplementation, secondary screening and transformation. Through knowledge graph and graph database technologies, design and generate a new structured model that conforms to the enterprise business;
[0009] S4: Perform model training and model evaluation on the new structured model:
[0010] After the evaluation criteria meet the enterprise requirements, the new structured model is applied to the enterprise's structured data model and participates in the screening of the collected structured and unstructured data;
[0011] After the evaluation criteria do not meet the enterprise requirements, the new structured model and the original data are discarded;
[0012] S5: The master data obtained from the structured data that meets the enterprise after ETL processing, the extracted metadata, and the enterprise business metadata screened in S2 are deeply integrated with AI technologies such as supervised learning, deep learning, regression models, and knowledge graphs and data quality management to achieve the evaluation of data cleaning and data quality. Finally, the evaluated data is input into the data resource location.
[0013] S6: The AI learning algorithm automatically identifies the usage frequency and popularity of data standards and inputs data through enterprise operations as the criteria for data quality evaluation, participating in the data quality evaluation in S5 to improve the level of data standard evaluation and the ability to optimize data.
[0014] Preferably, the data lake includes structured and unstructured data, and the AI data collection includes structured data collection and unstructured data collection.
[0015] Preferably, the data integration includes unstructured data integration and structured data integration. The methods of unstructured data integration and structured data integration include semantic models, classification and clustering algorithms, and automated data catalogs of label systems.
[0016] Preferably, the technical methods for the transformation and generation of the new structured model include knowledge graphs and graph database technologies.
[0017] Preferably, the learning process of the new structured model includes model training and model evaluation.
[0018] Preferably, the data quality evaluation includes the quality evaluation of master data, extracted metadata, and enterprise business metadata.
[0019] Preferably, the criteria for data quality evaluation include the usage frequency and popularity of data standards and data input through enterprise operations.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] This application aims to achieve a structured model for adaptive learning production through data lakes, data integration, etc. The data accessed by the data lake is automatically subjected to image recognition, text recognition, and speech recognition through AI technology. The AI data collects various structured and unstructured data and compares it with the enterprise's structured data model to screen out the structured data that conforms to the enterprise. The structured data that does not conform to the enterprise is integrated, and metadata in the integrated data is collected through an automated data catalog of semantic models, classification and clustering algorithms, and label systems. Then, data supplementation, secondary screening, and transformation are carried out, and through knowledge graph and graph database technologies, a new structured model that conforms to the enterprise's business is designed and generated. The new structured model undergoes model training and model evaluation. After the evaluation criteria meet the enterprise requirements, the new structured model is applied to the enterprise's structured data model to participate in the screening of the collected structured and unstructured data. If the evaluation criteria do not meet the enterprise requirements, the new structured model and the original data are discarded. In summary, a structured model for adaptive learning production is thus realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the data governance process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Based on the embodiments and drawings in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0024] As Figure 1 shown, the present invention provides a technical solution: an AI fusion governance method and process based on a data lake. The steps of this method are as follows:
[0025] S1: Connect the data lake data. By automatically performing image recognition, text recognition, and speech recognition on the data connected to the data lake through AI technology, various structured and unstructured data can be collected by AI data.
[0026] S2: Compare the collected structured and unstructured data with the enterprise's structured data model to screen out the structured data that conforms to the enterprise.
[0027] S3: Integrate the structured data that does not conform to the enterprise, collect metadata in the integrated data, then perform data supplementation, secondary screening, and transformation, and design and generate a new structured model that conforms to the enterprise's business through knowledge graph and graph database technologies.
[0028] S4: Perform model training and model evaluation on the new structured model:
[0029] After the evaluation criteria meet the enterprise requirements, apply the new structured model to the enterprise's structured data model and participate in the screening of the collected structured and unstructured data;
[0030] After the evaluation criteria do not meet the enterprise requirements, discard the new structured model and the original data;
[0031] S5: The master data obtained from the structured data that meets the enterprise requirements after ETL processing in S2, the extracted metadata, and the enterprise business metadata are deeply integrated with AI technologies such as supervised learning, deep learning, regression models, and knowledge graphs and data quality management to achieve the evaluation of data cleaning and data quality. Finally, the evaluated data is input into the data resource site.
[0032] S6: Automatically identify the usage frequency and popularity of data standards through AI learning algorithms and input data through enterprise operations as the criteria for data quality evaluation, and participate in the data quality evaluation in S5 to improve the level of data standard evaluation and the ability to optimize data.
[0033] Among them, the data lake includes structured and unstructured data, and the AI data collection includes structured data collection and unstructured data collection; due to the diversity of data in the data lake, key data required by enterprise operations are identified through AI technologies, and then structured data collection and unstructured data collection are carried out from the data lake to ensure the diversity and effectiveness of AI data collection.
[0034] Among them, the data integration includes unstructured data integration and structured data integration. The methods for unstructured data integration and structured data integration include semantic models, classification and clustering algorithms, and automated data catalogs of label systems; after comparing the unstructured data integration and structured data collected by AI data with the enterprise's structured data model, the unstructured data integration and structured data that do not meet the requirements are screened out, and the unstructured data integration and structured data that do not meet the requirements are integrated. During the integration process, the metadata in the unstructured data and structured data are mainly integrated through semantic models, classification and clustering algorithms, and automated data catalogs of label systems.
[0035] Among them, the technical methods for generating the new structured model include knowledge graphs and graph database technologies; after integrating the unstructured data integration and structured data that do not meet the requirements, data supplementation and screening are carried out, and the integrated data is designed into a more realistic enterprise business concept model through knowledge graphs and graph database technologies, and the concept model is transformed into a new structured model that can be recognized by the database and meets enterprise operations.
[0036] Among them, the new structured model learning process includes model training and model evaluation; after the new structured model passes model training and model evaluation and meets the enterprise requirements in terms of evaluation criteria, the new structured model is applied to the enterprise's structured data model to participate in the screening of the collected structured and unstructured data. If the evaluation criteria do not meet the enterprise requirements, the new structured model and the original data are discarded.
[0037] Among them, the data quality evaluation includes the quality evaluation of master data, extracted metadata, and enterprise business-oriented metadata. The criteria for the data quality evaluation include the usage frequency, popularity of data standards, and enterprise business input data; the usage frequency and popularity of data standards are automatically identified through AI learning algorithms and used as the criteria for data quality evaluation together with enterprise business input data, so as to improve the level of data standard evaluation and the ability to optimize data. Then, the master data, extracted metadata, and enterprise business-oriented metadata obtained from the ETL processing of the structured data that meets the enterprise requirements are deeply integrated with data quality management through AI technologies such as supervised learning, deep learning, regression models, and knowledge graphs to achieve data cleaning and data quality evaluation. Finally, the evaluated data is input into the data resource repository.
[0038] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0039] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI fusion governance method based on a data lake, characterized in that: The steps of this method are as follows: S1: Connect the data in the data lake. Automatically perform image recognition, text recognition, and speech recognition on the data connected to the data lake through AI technology, so as to collect various structured data and unstructured data through AI; S2: Compare the collected structured data and unstructured data with the enterprise's structured data model, and screen out the structured data that conforms to the enterprise; S3: Integrate the structured data that does not conform to the enterprise, collect metadata within the integrated data, then perform data supplementation, secondary screening and transformation, and design and generate a new structured model that conforms to the enterprise's business through knowledge graph and graph database technologies; S4: Perform model training and model evaluation on the new structured model: After the evaluation criteria meet the enterprise requirements, apply the new structured model to the enterprise's structured data model to participate in the screening of the collected structured data and unstructured data; After the evaluation criteria do not meet the enterprise requirements, discard the new structured model and the original data; S5: The master data obtained by ETL processing, the extracted metadata and the enterprise business metadata of the structured data that conforms to the enterprise screened in S2, through the deep integration of these AI technologies such as supervised learning, deep learning, regression model, knowledge graph and data quality management, realize the evaluation of data cleaning and data quality, and finally input the evaluated data into the data resource location; S6: Automatically identify the usage frequency and popularity of data standards through AI learning algorithms and input data through enterprise business, which are used as the criteria for data quality evaluation and participate in the data quality evaluation in S5 to improve the level of data standard evaluation and the ability to optimize data; Among them, the data integration includes unstructured data integration and structured data integration. The methods of unstructured data integration and structured data integration include semantic models, classification and clustering algorithms, and automated data catalogs of label systems; After comparing the unstructured data integration and structured data collected by AI with the enterprise's structured data model, screen out the unstructured data integration and structured data that do not conform, and perform data integration on the unstructured data integration and structured data that do not conform. During the integration process, integrate the metadata in the unstructured data and structured data through the automated data catalogs of semantic models, classification and clustering algorithms, and label systems; Among them, the technical methods for the transformation and generation of the new structured model include knowledge graph and graph database technologies; After integrating the unstructured data integration and structured data that do not conform, perform data supplementation and screening, and design a more realistic enterprise business concept model for the integrated data through knowledge graph and graph database technologies, and transform the concept model into a new structured model that can be recognized by the database and conforms to the enterprise's business.
2. The AI fusion governance method based on a data lake according to claim 1, wherein: The data lake includes structured data and unstructured data, and the AI data collection includes structured data collection and unstructured data collection.
3. The AI fusion governance method based on a data lake according to claim 1, characterized in that: The learning process of the new structured model includes model training and model evaluation.
4. The AI fusion governance method based on a data lake according to claim 1, characterized in that: The data quality assessment includes the quality assessment of master data, extracted metadata, and enterprise business-oriented metadata.
5. The AI fusion governance method based on a data lake according to claim 1, wherein: The criteria for the data quality assessment include the usage frequency, popularity of data standards, and enterprise business input data.
Citation Information
Patent Citations
Data processing method and device and data lake architecture
CN112597218A
Data handling methods and system for data lakes
US20180373781A1