Metadata management method and device based on data weaving architecture and medium
Through the metadata management method based on the data braiding architecture, metadata from multiple data sources is collected and analyzed, and knowledge graphs and recommendation engines are built, which solves the problems of chimneying of data pipelines and the surge in ETL tasks in big data management, and realizes efficient and intelligent data management and utilization.
Patent Information
- Application Number
- CN202510886241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the era of big data, enterprises are facing problems such as chimneying of data pipelines, surge in ETL tasks, high maintenance pressure from engineers, and high requirements for response efficiency and analysis flexibility in business scenarios.
Through a metadata management method based on the data weaving architecture, metadata from multiple data sources is collected, metadata warehouse and semantic relationships are determined, knowledge graphs are built, metadata mining and cataloging are established, recommendation engines and data service catalogs are established, and data service catalogs are realized efficient data management and utilization.
It realizes the comprehensiveness, intuitiveness and orderliness of metadata management, improves data semantic understanding capabilities, optimizes data model performance, improves data query efficiency and flexibility, supports the integration and display of multiple data sources, simplifies data integration analysis, and enhances data governance and security.
Smart Images

Figure CN120409651A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data governance, and particularly to a metadata management method, device, and medium based on a data weaving architecture. Background Art
[0002] With the advent of the big data era, data has become a new production factor, and its value has become increasingly prominent. The organization's demand for data is becoming more diverse, and massive amounts of data may be scattered in multiple application systems in a distributed environment. In particular, with the continuous accumulation of a large amount of semi-structured and unstructured data, the relevance of external data sources is increasing day by day, and the continuous development of the hybrid multi-cloud environment, the organization faces many challenges in the process of data management and application. For example, the current market application scenarios and participating teams are increasing, resulting in the "chimneyization" of data pipelines, and the workflows and data streams are complex; the data volume and ETL tasks are increasing rapidly, the maintenance pressure on engineers is high, and it is difficult to find optimization solutions; the business scenarios have higher requirements for response efficiency, analysis flexibility, etc. Summary of the Invention
[0003] To solve the above problems, this application proposes a metadata management method based on a data weaving architecture, including: collecting the metadata of the data source through a pre-set metadata management tool, and determining the corresponding metadata warehouse according to the metadata; determining the semantic relationship of the data source, and determining a knowledge graph according to the semantic relationship, so as to determine a semantic representation template according to the knowledge graph; performing mining processing on the metadata to determine a recommendation engine, and cataloging according to the metadata to determine a data service directory.
[0004] In one example, determining the corresponding metadata warehouse according to the metadata specifically includes: accessing multiple data sources through page configuration to delimit the scope of the multiple data sources, and then collecting data from the delimited multiple data sources, where the multiple data sources include structured data sources, semi-structured data sources, and unstructured data sources; determining the association relationship of the metadata, determining metadata assets according to the association relationship, processing the metadata assets, and determining the metadata warehouse according to the association relationship and the processed metadata assets; graphing the metadata warehouse and displaying the graphically represented metadata warehouse.
[0005] In one example, the method further includes: determining data distribution information, determining the metadata according to the data distribution information, determining the corresponding storage scheme for the metadata, and storing the metadata according to the storage scheme; determining data lineage according to the metadata warehouse, and spreading the metadata according to the data lineage.
[0006] In one example, after determining the knowledge graph according to the semantic relationship, the method further includes: determining the nodes and edges of the knowledge graph, representing the nodes as data information, and representing the edges as the relationships between the data information; quantifying the knowledge graph according to a preset algorithm to form a semantic knowledge graph.
[0007] In one example, the mining process of the metadata specifically includes: dividing the metadata to obtain multiple subsets, determining a training set and a test set according to the multiple subsets; determining a preset random forest, determining a feature subset according to the multiple subsets, splitting the feature subset, and recursively processing multiple sub-nodes of the random forest according to the split feature subset; determining the size of the random forest, traversing the random forest according to the size of the random forest, and training the random forest according to the training set.
[0008] In one example, the method further includes: determining multiple parameters of the random forest, setting the values of the multiple parameters according to a preset parameter range to obtain multiple parameter combinations; determining the combination effect of the multiple parameter combinations through cross-validation to determine the optimal parameters according to the combination effect.
[0009] In one example, the method further includes: determining a connector that supports multiple data sources, obtaining data information of the multiple data sources through the connector, and visually displaying the data information to determine a virtualization engine.
[0010] In one example, the method further includes: virtualizing the metadata to obtain a virtualized representation layer and data federation; providing a query service according to the virtualized representation layer, determining a query instruction, and querying a database according to the query instruction to obtain a query result.
[0011] On the other hand, the present application also proposes a metadata management device based on a data weaving architecture, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the metadata management device based on a data weaving architecture can execute: collecting metadata of a data source through a preset metadata management tool, and determining a corresponding metadata warehouse according to the metadata; determining the semantic relationship of the data source, determining a knowledge graph according to the semantic relationship, and determining a semantic representation template according to the knowledge graph; performing a mining process on the metadata to determine a recommendation engine, and cataloging according to the metadata to determine a data service catalog.
[0012] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to: collect metadata of a data source through a pre-set metadata management tool, and determine a corresponding metadata warehouse based on the metadata; determine the semantic relationship of the data source, and determine a knowledge graph based on the semantic relationship, so as to determine a semantic representation template based on the knowledge graph; mine the metadata to determine a recommendation engine, and catalog the metadata to determine a data service directory.
[0013] This application uses metadata management tools to collect metadata and determine the corresponding warehouse, supports access to multiple data sources, delimits the scope of data collection, and can also determine the association relationship to form metadata assets and process them, and finally display them in a graphical manner, making metadata management more comprehensive, intuitive and orderly. Determine the data distribution information and storage plan, and spread metadata based on data lineage, which is conducive to the reasonable storage and efficient use of data. After determining the knowledge graph, it is quantified to form a semantic knowledge graph, which can clearly represent data information and relationships and improve the ability to understand data semantics. When mining metadata, dividing the data set, determining feature subsets, and recursively processing, traversing and training the random forest will help to tap the potential value of the data. Determining multiple parameter combinations and determining the optimal parameters through cross-validation can optimize model performance. Determine a connector that supports multiple data sources to obtain data and visualize it, and determine a virtualization engine to facilitate data integration and display. Metadata virtualization processing obtains a virtualized representation layer and data federation, provides query services, can quickly respond to query instructions to obtain results, and improve data query efficiency and flexibility. Overall, this method forms a complete, efficient and intelligent metadata management system from data collection, storage, association, mining, visualization to query services, which helps enterprises better manage and utilize data resources and enhance data value. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 Schematic diagram of a metadata management method based on a data weaving architecture in an embodiment of the present application; Figure 2 This is a schematic diagram of a metadata management device based on a data weaving architecture in an embodiment of the present application. DETAILED DESCRIPTION
[0015] To make the objectives, technical solutions and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0016] The following details the technical solutions provided by each embodiment of this application in conjunction with the drawings.
[0017] As Figure 1 shown, to solve the above problems, a metadata management method based on a data weaving architecture provided by an embodiment of this application includes: S101. Collect the metadata of the data source through a pre-set metadata management tool, and determine the corresponding metadata warehouse according to the metadata.
[0018] When constructing an active metadata management tool, first establish connections with relational databases, big data platforms, data warehouses, document storage systems, APIs, Web services, etc. Scan and parse various data sources of the enterprise. For the accessible data source types, including structured data sources, semi-structured data sources, and unstructured data sources, provide a page configuration method to configure the data source information and delimit the scope of the accessed data resources. Structured data sources include relational databases, semi-structured data sources include relational databases, and unstructured data sources include distributed file systems, FTP data, etc.
[0019] Utilize an automated pipeline to configure the metadata acquisition time, acquisition object, and processing script, and regularly acquire all metadata information related to data according to the configuration. Based on the result information obtained in the above steps, construct the relationship mapping between tables and fields to form metadata assets. According to the existing business relationship dictionary table, process the metadata assets, establish the association relationship between tables, fields, and business, and manually check the processing results to finally form a metadata warehouse. Finally, through the metadata automatic acquisition technology, timely query the changed metadata of external data sources, store it in the metadata warehouse, and automatically update the metadata to generate a metadata comparison record.
[0020] In one embodiment, for data such as database tables, documents, pictures, audio, and video generated by a business system, a page configuration function is provided to configure data source information, clarify the scope of data resources to be accessed, and collect metadata information of relevant data resources. The types of data sources that can be accessed cover structured data sources, semi-structured data sources, and unstructured data sources. Among them, structured data sources and semi-structured data sources mainly include databases; unstructured data sources include distributed file systems and FTPs. By configuring the information of heterogeneous data sources, the corresponding data sources can be connected, and the data access scope can be set in the data sources. Specifically, if the data source type is a structured data source or a semi-structured data source, the range of database tables to be accessed needs to be set in the data source, and the metadata information of the tables is collected, including but not limited to table names, table remarks, field names, field types, field remarks, primary and foreign key information, and table connection information, etc.; if the data source type is an unstructured data source, the documents, pictures, audio, and video data to be accessed need to be set in the distributed file system or FTP, and information such as the file name, creation time, creator, file size, and storage location is collected.
[0021] In an automated pipeline manner, configure the acquisition time, acquisition object, and processing script of the metadata, and regularly acquire all metadata information related to the data according to the configuration. For the accessed metadata, there are three cases for the associated calculation methods: library table resources and library table resources, library table resources and file resources, and file resources and file resources. Automatically calculate the metadata association relationships of different data resources through a text similarity algorithm, and then form a metadata asset. Process the metadata asset, establish the association relationships between tables, fields, and the business, and manually check the processing results to finally form a metadata warehouse. In addition, the data in the metadata warehouse can also be graphically displayed, and the blood relationship between the metadata can be optimized and adjusted through manual intervention.
[0022] S102. Determine the semantic relationship of the data source, determine a knowledge graph according to the semantic relationship, and determine a semantic representation template according to the knowledge graph.
[0023] Dynamic knowledge graph modeling constructs a metadata knowledge graph based on the semantic relationships of multi-source data to generate a unified semantic representation. Using the Neo4j component, the data in the metadata warehouse is graphically displayed, and the lineage relationship between metadata can be optimized through manual intervention. Active metadata is compared with static passive metadata. It not only defines the data itself but also covers all the operations occurring on the data and the data generated during the process, mainly including three key components: the metadata warehouse, data process automation, and knowledge graph construction. Among them, the metadata warehouse, as the cornerstone of active metadata, contains all data related to data and generated by operating on data, such as technical metadata, business metadata, operation metadata, social metadata, etc. This helps to enrich the business semantics of the data and enables the system and data consumers to better understand the data.
[0024] Data flow automation includes automatically collecting data distribution information, automatically classifying sensitive data and business data across the entire domain, and simultaneously performing real-time diffusion based on data lineage to achieve classified and hierarchical data management and compliance governance strategies. Through the real-time collection and parsing of full-link SQL logs, operator-level data lineage can be automatically parsed and generated, and intuitive field processing specifications can be extracted. Users can quickly conduct impact analysis, trace the source, track downstream, and identify key nodes with the help of the lineage visualization UI, thereby efficiently carrying out data governance analysis work. In the process of constructing a knowledge graph, in addition to preprocessing the data in the metadata warehouse, entity recognition and linking are key steps. Entity recognition is to identify entities with unique identifiers from the metadata warehouse, such as people, places, organizations, etc.; entity linking is to link these entities to the corresponding entities in the existing knowledge base to construct a knowledge graph. The technologies of entity recognition and linking cover traditional rule- and dictionary-based methods, as well as machine learning- and deep learning-based methods. These technologies can improve the accuracy and efficiency of entity recognition and linking by identifying the context information of entities, named entity recognition, entity disambiguation, etc. Relationship extraction in knowledge graph construction is to extract the semantic relationships between entities from text for constructing the connections between entities in the knowledge graph. Relationship extraction technologies include rule-based pattern matching, machine learning-based relationship extraction models, and deep learning-based relationship extraction methods, etc. These technologies can achieve the accuracy and robustness of relationship extraction by mining the text features, semantic information, and syntactic structures between entities. In addition, a knowledge graph does not represent data information in rows, columns, tables, and keys, but uses nodes and edges to represent data assets and the relationships between them. Fundamentally speaking, this graphical data model is simpler than the relational model, yet more expressive and functional, easier to modify and infinitely extensible. Moreover, the knowledge graph actually exists in the computing layer of the data management system, rather than the storage layer, which means that it can be modified at any time by adding new nodes and edges, without having to painstakingly conceive a single shared data model that covers all current and future organizational data needs at a certain point in time.
[0025] Furthermore, through AI / ML algorithms for entity connection and quantification of connection relationships, the association relationships between "data and data, data and users, data and business semantics" are automatically mined and established to form a semantic knowledge graph, thereby achieving a more three-dimensional characterization of data and supporting more intelligent data usage recommendations.
[0026] In one embodiment, active metadata covers two major aspects: the metadata warehouse and data process automation. The metadata warehouse is the foundation and core of active metadata. It not only contains technical metadata, business metadata, operational metadata, social metadata, but also includes all operation records related to data and the data generated by all activities carried out on the data. Data process automation includes a series of key functions: automatically collecting data distribution information, automatically capturing and generating metadata; selecting appropriate storage solutions according to data characteristics and requirements; establishing real-time or regular update mechanisms to ensure the timeliness and accuracy of metadata. In addition, process automation can automatically classify sensitive data and business data across the board, and perform real-time dissemination based on data lineage, so as to achieve classified and hierarchical management of data, and formulate and implement compliant governance strategies.
[0027] In one embodiment, the knowledge graph uses nodes and edges to represent data information and the association relationships between this information. With the help of artificial intelligence or machine learning algorithms, entities are connected and the connection relationships are quantitatively analyzed, so as to automatically mine and construct the internal connections between data and data, data and users, and data and business semantics, and finally form a semantic knowledge graph.
[0028] In one embodiment, the machine learning algorithm of random forest is used to carry out model training work on the metadata information in the metadata warehouse. Use the random forest algorithm to complete operations such as training and business fitting. Based on the result set formed by this algorithm, match it with data users and sort it according to the matching degree.
[0029] S103. Mine and process the metadata to determine a recommendation engine, and catalog according to the metadata to determine a data service catalog.
[0030] Based on business experience and machine learning models, deeply mine metadata to build an intelligent recommendation engine. This recommendation engine integrates the rules formed by expert experience and machine learning models and is applied to data management, data preparation, and services, such as formulating data integration solutions and optimizing engine performance. Its recommendation scope covers all stages of the data full life cycle, including data asset recommendation, data usage recommendation, data integration solution recommendation, execution plan recommendation, computing engine recommendation, data classification suggestion, data timeliness improvement suggestion, data security risk control suggestion, cost governance suggestion, etc. Among them, the machine learning model uses the random forest algorithm to train the model and fit the business for the metadata information in the metadata warehouse. Based on the result set formed by the algorithm, it matches with data users and sorts according to the matching degree, and recommends relevant data tables to business personnel, reducing the time for them to find data. Truly effective data weaving is intelligent. The system needs to provide automatic suggestions according to the specific use of data, combined with task load and enterprise data management requirements, and be able to analyze past activities to predict the future, so as to form the most appropriate recommendation suggestions. Therefore, the recommendation engine needs to have good openness to support the implementation of various recommendation schemes and stable services. The functions of the intelligent recommendation engine cover intelligent data classification and intelligent SQL association. Intelligent data classification is based on active metadata and field content sampling, automatically identifies PII sensitive information, recommends the business classification of data assets, and spreads classification labels in real time based on operator-level lineage, completing global data classification with low cost, high timeliness, and high accuracy, providing basic data support for data classification and grading management. Intelligent SQL association, on the other hand, identifies senior data users and deeply mines their data usage behaviors. When users write SQL code, in addition to providing SQL syntax hints, it can also associate data usage, such as common table associations, common filtering conditions, common aggregation dimensions, and measurement fields, greatly improving the efficiency of users' SQL writing and the experience of using data.
[0031] With active metadata at the core, leveraging AI artificial intelligence and machine learning technologies to achieve automated data cataloging, and then constructing an enhanced data service catalog. When analysts and business personnel perform self-service, the primary challenge is how to discover, understand, and trust data from vast and scattered data. Different from traditional data dictionary or data map products, the enhanced data catalog aims to be accessible to all users within the enterprise through an intuitive user interface, not just IT personnel. Since analysts and business personnel are not familiar with the enterprise's data asset situation, when using data dictionary or data map products to search for data self-service, they often face problems such as difficulty in finding data, cumbersome operation, and hesitation to use. The emergence of the enhanced data catalog aims to enable analysts to quickly search and locate data, evaluate and select the most suitable data, and thus carry out data preparation and analysis work efficiently and confidently. Further, with active metadata at the core, applying AI and machine learning to metadata collection, semantic reasoning, and classification tagging to achieve automatic data cataloging, minimizing the manual maintenance of metadata work to the greatest extent. The enhanced data service catalog can provide the following services to business personnel: First, semantic data search, providing a powerful search function friendly to business personnel, supporting keyword, business term, and natural language searches, and sorting search results according to relevance and usage frequency to help users quickly locate the required data; Second, panoramic data profiling, by comprehensively and deeply depicting data, helping users evaluate the matching degree between data and analysis requirements, such as providing data sampling previews, quality information, output timeliness, security sensitivity levels, user ratings and evaluations, expert annotations, common usages, etc. These information are automatically generated by active metadata, which can significantly improve the efficiency of users in selecting data; Third, visual lineage analysis, providing users with an intuitive lineage analysis tool, facilitating users to flexibly explore the upstream and downstream links of data, intelligently discover key nodes and paths, and quickly clarify the data context; Fourth, global data search, providing the technical ability to directly access global data for interactive federated queries, while built-in access protection mechanisms for sensitive data such as security, privacy, and compliance.
[0032] In one embodiment, when performing in-depth mining on metadata, dataset partitioning is first carried out. The original dataset is divided into a training set and a test set, usually using the K-fold cross-validation method. Specifically, the dataset is evenly divided into K subsets. Each time, one of the subsets is selected as the test set, and the remaining K - 1 subsets are used as the training set. This process is repeated K times, and finally the average value is taken as the performance evaluation index of the model. Next is to construct decision trees. In a random forest, each decision tree is constructed independently. When constructing each decision tree, a feature subset is randomly selected from the original dataset. The size of the feature subset is generally set to 1 / 3 or 1 / 4 of the total number of features, and then splitting is performed based on this feature subset. Then the same operation is recursively performed on each child node until the stopping condition is met, such as the number of samples in the node being less than the threshold, reaching the maximum depth, etc. Then comes ensemble learning. In a random forest, the prediction results of all decision trees are used to calculate the final classification or regression result. For classification problems, a voting method is adopted. Each decision tree votes according to its own prediction result, and the class with the most votes is taken as the final result. For regression problems, the average value of the prediction results of all decision trees is taken as the final result. Finally is the training process. The specific steps are as follows: set the size of the training set T to N, the number of features to M, and the size of the random forest to K; traverse the size of the random forest K times; sample N times from the training set T with replacement to form a new sub-training set D; randomly select m features, where m < M; use the new training set D and these m features to learn a complete decision tree; repeat the above process to finally obtain a random forest.
[0033] In one embodiment, a data virtualization engine is constructed through technologies such as federated query, dynamic integration, and data orchestration to achieve full-link self-service data usage. The core of the federated query is based on Apache Hive3 and SQL to implement cross-database queries. Apache Hive automatically identifies the data sources in the query statement to be executed according to the configured JDBC data sources, realizes intelligent JDBC pushdown with the help of a cost-based optimizer, automatically groups the data sources in the query statement, and finally generates a result set that matches the query statement. Data virtualization is the key to achieving data weaving, and it undertakes the important responsibility of enabling business users to complete data integration, preparation, and delivery by themselves. It constructs a virtual semantic layer between the data source and the data consumption end for connecting, integrating, and consuming data. Users can complete data transformation by defining data queries, thereby achieving transparent integration, self-service preparation, and high-performance services for cross-source and cross-environment data. The specific steps for constructing a data virtualization engine include: creating connectors that support different data sources, such as JDBC, ODBC, MQTT, AMQP, etc.; obtaining information such as tables and fields in the target data source (or database) through the connectors and intuitively displaying the data in the data source in a front-end visualization manner; creating data services for adding, deleting, querying, and modifying object tables in the form of automated APIs (application programming interfaces); providing a visualization interface through which users can combine and orchestrate data services from multiple different databases by dragging, pulling, and dropping to form data services related to the business and provide them for external calls in the form of API interfaces.
[0034] In one embodiment, data virtualization generally includes two core components: the data virtualization presentation layer and data federation. The data virtualization presentation layer can provide query services at the virtual layer or semantic layer, shielding the storage details of the underlying database and making the data appear as a single data model. When the underlying data federation mechanism receives a query request, it decomposes it into a query part for a relational database, executes the actual data query operation, and returns the result. The whole process not only avoids a large amount of data migration and replication work but also provides a unified data application view, making the formatting and management details of the data in its original source transparent to data consumers. Finally, it enables consumers to define the form of data return and combine data from multiple sources according to this form.
[0035] In one embodiment, the method of grid search is used to optimize the parameters of the random forest algorithm. Specifically, each parameter to be optimized is set to multiple values within a range, and then the effects of these parameter combinations are evaluated through cross-validation. During this process, key parameters such as the number of trees (n_estimators), maximum depth (max_depth), and minimum number of samples per leaf (min_samples_leaf) are focused on. To further optimize the random forest model, the idea of ensemble learning is also introduced. By constructing multiple independent random forest models and combining their results, the generalization ability and robustness of the model can be significantly improved. The method adopted here belongs to the Bagging idea in ensemble learning. By adjusting the weights of each sub-model in the final result, the trade-off of different feature importances is realized, so as to find the optimal solution.
[0036] In one embodiment, the federated query is a cross-database query method implemented based on Apache Hive3 and the SQL structured query language. Apache Hive3 will automatically identify the data sources involved in the query statement according to the configured JDBC data source connection. It uses a cost-based optimizer to implement the JDBC intelligent pushdown function, automatically group the data sources in the query statement, and finally generate a result set that matches the query statement.
[0037] In one embodiment, when building a data virtualization engine. First, create connectors that can support different data sources; then, use these connectors to obtain information such as tables and fields in the target data source or database, and visually display the data in the data source through the front-end visualization interface; then, create data services such as adding, deleting, querying, and modifying based on object tables in the form of an automated API application interface; finally, provide a visualization operation interface, and users can combine and orchestrate the data services of multiple different databases through convenient operations such as dragging, pulling, and dropping to form data services closely related to the business, and provide a call function externally in the form of an API interface.
[0038] In one embodiment, data virtualization includes two key parts: the data virtualization presentation layer and data federation. Among them, the data virtualization presentation layer provides query services for users at the virtual layer or semantic layer. Its function is to shield the storage details of the underlying database, so that users can perform query operations without paying attention to the specific storage method of the underlying database. After receiving the query instruction, the data federation mechanism will decompose the instruction into query sub-parts for different databases, then execute these sub-queries respectively to perform the actual data query operation, and finally return the query result to the user.
[0039] In one embodiment, data virtualization-related initiatives can bring significant benefits to enterprises in many aspects. First, it enhances the data user experience and accelerates data delivery. Enterprises establish a global data catalog and apply AI technologies such as semantic search, knowledge graphs, and NLP, enabling users to conveniently and quickly obtain rich, trustworthy, and high-quality data. In this way, users can focus more time on business scenarios and data analysis rather than spending it on searching for and identifying data. Second, it simplifies the integrated analysis mode and solves the data silo problem. Enterprises use virtual data link methods to organically connect scattered, dynamic, and diverse data sources, breaking down the barriers to data integration and correlation analysis. The data remains in its original location, and exploration and access can be carried out without developing and deploying ETL jobs. As new data sources are quickly integrated into the system and the full-process intelligence continues to evolve, the scale and user experience of the data weaving system will continue to improve, effectively avoiding access restrictions caused by data stored in different environments and effectively eliminating data silos. At the same time, the change in the integration method significantly reduces the same data copies, thereby reducing the costs of data storage, maintenance, and management. Third, it helps with comprehensive data governance and strengthens security and privacy protection. Data weaving helps enterprises achieve comprehensive data governance, formulate unified access control and privacy protection strategies, ensure controllable risks during the data analysis and application process, meet regulatory requirements, and prevent the leakage of privacy data. Fourth, it gains insights into users' data needs and builds a smart data usage community. By recording users' data access trajectories and analyzing their data usage, combination preferences, and access patterns, on the one hand, it can discover more business application cooperation and sharing opportunities and promote the deeper application of data; on the other hand, it can proactively recommend or automatically optimize data distribution and flow, promote the precipitation of common data, reduce the frequent large-scale cross-regional and cross-system data flows, improve users' access efficiency, and reduce the overall hardware resource overhead.
[0040] As Figure 2 shown, the embodiment of the present application also provides a metadata management device based on a data weaving architecture, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the metadata management device based on a data weaving architecture can execute: Collect metadata of a data source through a pre-set metadata management tool, and determine a corresponding metadata warehouse according to the metadata; Determine the semantic relationship of the data source, determine a knowledge graph according to the semantic relationship, and determine a semantic representation template according to the knowledge graph; Mine the metadata to determine a recommendation engine, and catalog according to the metadata to determine a data service catalog.
[0041] An embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, where the computer-executable instructions are set as: Collect metadata of a data source through a pre-set metadata management tool, and determine a corresponding metadata warehouse according to the metadata; Determine the semantic relationship of the data source, determine a knowledge graph according to the semantic relationship, and determine a semantic representation template according to the knowledge graph; Mine the metadata to determine a recommendation engine, and catalog according to the metadata to determine a data service catalog.
[0042] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0043] The device and medium provided by the embodiments of the present application correspond one by one to the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.
[0044] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0045] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing in the process Figure 1One or more processes and / or blocks Figure 1 Apparatus for the functions specified in one or more blocks
[0046] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the processes Figure 1 One or more processes and / or blocks Figure 1 The functions specified in one or more blocks
[0047] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the processes Figure 1 One or more processes and / or blocks Figure 1 Steps for the functions specified in one or more blocks
[0048] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory
[0049] Memory may include non-permanent memory in computer-readable media, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media
[0050] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves
[0051] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0052] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A metadata management method based on a data weaving architecture, characterized in that, Including: Collecting metadata of a data source through a pre-set metadata management tool, and determining a corresponding metadata warehouse according to the metadata; Determining the semantic relationship of the data source, determining a knowledge graph according to the semantic relationship, and determining a semantic representation template according to the knowledge graph; Mining and processing the metadata to determine a recommendation engine, and cataloging according to the metadata to determine a data service catalog.
2. The method according to claim 1, characterized in that Determining a corresponding metadata warehouse according to the metadata, specifically including: Accessing multiple data sources through page configuration to delimit the scope of the multiple data sources, and then collecting data from the delimited multiple data sources, where the multiple data sources include structured data sources, semi-structured data sources, and unstructured data sources; Determining the association relationship of the metadata, determining metadata assets according to the association relationship, processing the metadata assets, and determining the metadata warehouse according to the association relationship and the processed metadata assets; Graphing the metadata warehouse and displaying the graphically represented metadata warehouse.
3. The method according to claim 1, characterized in that, The method further includes: Determining data distribution information, determining the metadata according to the data distribution information, determining a storage scheme corresponding to the metadata, and storing the metadata according to the storage scheme; Determining data lineage according to the metadata warehouse and spreading the metadata according to the data lineage.
4. The method according to claim 1, wherein After determining the knowledge graph according to the semantic relationship, the method further includes: Determining the nodes and edges of the knowledge graph, representing the nodes as data information, and representing the edges as the relationships between the data information; Quantifying the knowledge graph according to a pre-set algorithm to form a semantic knowledge graph.
5. The method according to claim 1, characterized in that, Mining and processing the metadata, specifically including: Dividing the metadata to obtain multiple subsets, and determining a training set and a test set according to the multiple subsets; Determining a pre-set random forest, determining a feature subset according to the multiple subsets, splitting the feature subset, and recursively processing multiple child nodes of the random forest according to the split feature subset; Determining the size of the random forest, traversing the random forest according to the size of the random forest, and training the random forest according to the training set.
6. The method according to claim 5, characterized in that, The method further includes: Determining multiple parameters of the random forest, setting values of the multiple parameters according to a pre-determined parameter range to obtain multiple parameter combinations; Determining the combined effect of the multiple parameter combinations through cross-validation to determine optimal parameters according to the combined effect.
7. The method according to claim 1, characterized in that The method further includes: Determining a connector that supports multiple data sources, obtaining data information of the multiple data sources through the connector, and visually displaying the data information to determine a virtualization engine.
8. The method according to claim 1, wherein The method further includes: Virtualizing the metadata to obtain a virtualized presentation layer and a data federation; Providing a query service according to the virtualized presentation layer, determining a query instruction, and querying a database according to the query instruction to obtain a query result.
9. A metadata management device based on a data weaving architecture, characterized in that, Including: At least one processor; And, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the metadata management device based on the data weaving architecture can execute: collect metadata of a data source through a pre-set metadata management tool, and determine a corresponding metadata warehouse according to the metadata; determine the semantic relationship of the data source, determine a knowledge graph according to the semantic relationship, so as to determine a semantic representation template according to the knowledge graph; perform mining processing on the metadata to determine a recommendation engine, and perform cataloging according to the metadata to determine a data service catalog.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are configured to: collect metadata of a data source through a pre-set metadata management tool, and determine a corresponding metadata warehouse according to the metadata; determine the semantic relationship of the data source, determine a knowledge graph according to the semantic relationship, so as to determine a semantic representation template according to the knowledge graph; perform mining processing on the metadata to determine a recommendation engine, and perform cataloging according to the metadata to determine a data service catalog.
Citation Information
Patent Citations
Metadata graph-based metadata organization management method and system
CN107291875A
Data management method based on data weaving architecture
CN116303336A
Cited By
Data governance strategy dynamic execution system based on active metadata and AI recommendation
CN121478755A