A unified management method and system for multi-source heterogeneous data based on metadata

By using a unified management approach based on metadata, multi-source heterogeneous data is characterized, classified, and managed in layers. This solves the problem of the lack of integration of multi-source heterogeneous data, realizes standardized and visualized data management, and improves data utilization efficiency and the clarity of enterprise data assets.

CN115729993BActive Publication Date: 2026-03-10CHINA COMM SERVICE APPL & SOLUTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In existing technologies, multi-source heterogeneous data cannot be fully integrated, resulting in low data utilization efficiency, inability to adapt to flexible expansion of business needs, and a lack of effective control over data management objects.

Method used

A unified management approach based on metadata is adopted, which describes and manages multi-source heterogeneous data through preset definition rules. Metadata definition methods are provided for both new and existing data to form standardized data management. Data is also classified and hierarchically managed through a data catalog to achieve data visualization and unified management.

Benefits of technology

It enables standardized management of multi-source heterogeneous data, improves data utilization efficiency, meets users' needs for data cognition and retrieval, forms clear data assets for enterprises, supports enterprises in quickly understanding the overall situation of data assets, and enhances the support for data applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115729993B_ABST
    Figure CN115729993B_ABST
Patent Text Reader

Abstract

This invention discloses a unified management method and system for multi-source heterogeneous data based on metadata. The method includes: responding to data producer users' upload of metadata for multi-source heterogeneous data, and converting the metadata to a runtime state; wherein, the metadata is data obtained by data producer users based on preset metadata definition rules to characterize multi-source heterogeneous data; configuring corresponding permissions for users with different roles; responding to platform management users and / or data producer users' data catalog compilation operations, establishing associations between data catalogs and data objects based on the compilation operations, forming a categorized and hierarchical data catalog; and returning corresponding data attribute information and content information to data consumer users. This invention, through unified metadata characterization and data catalog encapsulation of multi-source heterogeneous data, forms standardized and visualized data management, meeting users' needs for data cognition and data retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data management technology, specifically relating to a unified management method and system for multi-source heterogeneous data based on metadata. Background Technology

[0002] Currently, common data categories include relational data, file data, and message data. Relational data refers to data managed through relational databases and displayed as a two-dimensional logical table structure. File data refers to data stored as file objects in a specific format. Message data refers to data produced, exchanged, and used through message processing middleware such as MQ / Kafka. Currently, the use of these three types of heterogeneous data is not fully integrated. Data processing capabilities are usually limited to the same category of data, and analysis and comparison between different categories often rely on manual methods. For example, transmitting data content through pre-agreed relational database middleware tables cannot adapt to flexible expansion needs and results in low data utilization efficiency. For file and relational data, whether storing file objects directly in a predetermined server location or using mainstream message middleware to subscribe to or consume message objects without local storage of message content, it is merely a black-box management approach, failing to effectively grasp the specific information of the data management objects.

[0003] Therefore, there is an urgent need to propose a method for unified management of multi-source heterogeneous data to solve the above problems. Summary of the Invention

[0004] The purpose of this invention is to provide a unified management method and system for multi-source heterogeneous data based on metadata, in order to solve the technical problems of insufficient integration and low data utilization efficiency in the existing technology of multi-source heterogeneous data.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] The first aspect provides a unified management method for multi-source heterogeneous data based on metadata, including:

[0007] The system responds to data producers uploading metadata of multi-source heterogeneous data, receives, reviews, and verifies the metadata, and converts the metadata to the runtime state after successful verification. The metadata is data obtained by data producers based on preset metadata definition rules to characterize multi-source heterogeneous data, and the multi-source heterogeneous data includes at least relational data, file data, or message data.

[0008] Configure corresponding permissions for users with different roles so that each user can perform corresponding operations on the data directory based on their own permissions; among them, users with different roles include data producers, platform managers, and data consumers;

[0009] Responding to the data catalog compilation operations of platform management users and / or data production users, establishing the association between the data catalog and data objects based on the compilation operations, forming a classified and hierarchical data catalog; wherein, the data objects are instantiated metadata objects;

[0010] In response to data consumer users' information retrieval operations on the data catalog, the system returns the corresponding data attribute information and content information to the data consumer users.

[0011] In one possible design, the metadata includes new data and existing data;

[0012] When the metadata is newly added data, before responding to the data producer user's operation of uploading the metadata of multi-source heterogeneous data, the method further includes: receiving the data producer user's data production request, generating a data demand form and associating it with the corresponding batch number;

[0013] When the metadata is existing data, before responding to the data producer user's operation of uploading metadata of multi-source heterogeneous data, the method further includes:

[0014] Analyze the existing data to be characterized, including basic information and content elements, to form a characterization template and publish it.

[0015] In one possible design, the preset metadata definition rules define metadata based at least on the dimensions of category, location, content, and parsing parameters; wherein, the definition of parsing parameters adopts a predefined approach to constrain the extended definition of content elements.

[0016] In one possible design, the metadata is received, reviewed, and verified, and after successful verification, the metadata is converted to the runtime state, including:

[0017] The metadata is received, and the compliance and consistency of the basic information and content elements defined in the metadata are reviewed. After the review is passed, the version release verification is initiated.

[0018] The metadata within the version is archived using the version number, and the metadata is then converted to the runtime state; wherein, the version number includes at least one batch number;

[0019] When the metadata is existing data, after archiving the metadata within a version using the version number, it also includes:

[0020] Update the metadata design table and the release history table.

[0021] In one possible design, corresponding permissions are configured for users with different roles, so that each user can perform corresponding operations on the data directory based on their own permissions, including:

[0022] Configure platform management users with at least the permissions for directory configuration management, directory classification, directory approval, and directory extension approval, so that platform management users can perform data directory planning, top-level directory creation and maintenance, data directory approval, and data directory extension approval based on their own permissions;

[0023] Configure at least the cataloging permissions for data production users so that they can expand and compile data catalogs based on their own permissions;

[0024] Configure at least directory retrieval permissions for data consumer users so that they can query data directory information based on their own permissions.

[0025] In one possible design, data producers extend and compile the structure of the data directory and the mounting of data objects based on preset classification rules, classification abbreviations, and directory encoding rules.

[0026] In one possible design, the association between the data catalog and data objects is established based on the compilation operation, including:

[0027] Obtain the tag information of data objects under the data directory described by the data producer user, and establish the association between the data directory and the data objects under the corresponding data source based on the tag information.

[0028] In one possible design, after forming a categorized and hierarchical data catalog, the method further includes:

[0029] Publish the completed data catalog and set the access method, scope, and time for the data catalog.

[0030] In one possible design, responding to a data consumer's information retrieval operation on the data catalog includes:

[0031] A unified query window is used to respond to information retrieval operations by data consumers, allowing them to query information based on the data directory structure and data objects.

[0032] The second aspect provides a unified management system for multi-source heterogeneous data based on metadata, including:

[0033] The metadata publishing module is used to respond to the operation of data producer users uploading metadata of multi-source heterogeneous data, receive, review and verify the metadata, and convert the metadata to the running state after the verification is passed; wherein, the metadata is the data obtained by data producer users based on preset metadata definition rules to characterize multi-source heterogeneous data, and the multi-source heterogeneous data includes at least relational data, file data or message data;

[0034] The permission configuration module is used to configure corresponding permissions for users with different roles, so that users with different roles can perform corresponding operations on the data directory based on their own permissions; among them, users with different roles include data producers, platform managers, and data consumers.

[0035] The data catalog publishing module is used to respond to the data catalog compilation operations of platform management users and / or data production users. Based on the compilation operations, it establishes the association between the data catalog and data objects and data source storage information, forming a classified and hierarchical data catalog; among them, data objects are instantiated metadata objects;

[0036] The data query module is used to respond to data consumer users' information retrieval operations on the data catalog and return the corresponding data attribute information and content information to the data consumer users.

[0037] Thirdly, the present invention provides a computer device comprising a memory, a processor, and a transceiver connected in sequence and communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the metadata-based multi-source heterogeneous data unified management method as described in any possible design of the first aspect.

[0038] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform a metadata-based unified management method for heterogeneous multi-source data as described in any possible design of the first aspect.

[0039] Fifthly, the present invention provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform a metadata-based unified management method for heterogeneous multi-source data as described in any possible design of the first aspect.

[0040] The advantages of this invention compared to the prior art are:

[0041] This invention, based on preset metadata definition rules, provides unified metadata characterization and management for multi-source heterogeneous data, applicable to business data in various formats from multiple file-based and / or message-based data sources. By providing corresponding metadata definition methods for new and existing data, it defines data characteristics to form standardized data management. Based on metadataization and data catalog abstraction and encapsulation methods, it forms categorized and hierarchical data assets, enabling visualized data management and meeting users' needs for data cognition and retrieval. After standardizing the metadata management of file-based or message-based data, and then classifying and hierarchically managing it through a data catalog, it forms clear data assets for enterprises, allowing them to quickly and clearly understand the overall status of their data assets and providing strong data support for subsequent data applications. Attached Figure Description

[0042] Figure 1 This is a flowchart of a method for unified management of multi-source heterogeneous data based on metadata, as described in the embodiments of this application.

[0043] Figure 2 This is a logical diagram illustrating the unified management method for multi-source heterogeneous data based on metadata in the embodiments of this application;

[0044] Figure 3 This is a schematic diagram illustrating the metadata characterization of newly added data in an embodiment of this application;

[0045] Figure 4 This is a logical diagram illustrating the metadata characterization of existing data in an embodiment of this application;

[0046] Figure 5 This is a flowchart illustrating the data directory abstraction and encapsulation process in an embodiment of this application;

[0047] Figure 6 This is a flowchart illustrating the data catalog management process in an embodiment of this application;

[0048] Figure 7 This is a schematic diagram of the data catalog classification model in the embodiments of this application. Detailed Implementation

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0050] Example

[0051] To address the technical problem of insufficient integration and low data utilization efficiency in existing technologies involving multi-source heterogeneous data, this application provides a metadata-based unified management method for multi-source heterogeneous data. This method uses preset metadata definition rules to uniformly characterize and manage multi-source heterogeneous data, applicable to business data in various formats from multiple file-based and / or message-based data sources. It defines data characteristics by providing corresponding metadata definition methods for new and existing data, forming standardized data management. Based on metadataization and data catalog abstraction and encapsulation methods, it forms categorized and hierarchical data assets, enabling visualized data management and meeting users' needs for data understanding and retrieval. After standardizing the metadata management of file-based or message-based data, and then classifying and hierarchically managing it through a data catalog, it forms clear data assets for the enterprise, allowing the enterprise to quickly and clearly understand the overall status of its data assets and providing strong data support for subsequent data applications.

[0052] The following detailed description of the metadata-based unified management method for heterogeneous multi-source data in this application is provided through specific embodiments.

[0053] like Figures 1-7 As shown, this application provides a method for unified management of multi-source heterogeneous data based on metadata, including but not limited to steps S1-S4:

[0054] Step S1. Respond to the data producer user's operation of uploading metadata of multi-source heterogeneous data, receive, review and verify the metadata, and convert the metadata to the running state after verification; wherein, the metadata is the data obtained by the data producer user based on preset metadata definition rules to characterize multi-source heterogeneous data, and the multi-source heterogeneous data includes at least relational data, file data or message data;

[0055] It should be noted that, compared to the traditional passive and inflexible approach of metadata model characterization, this application provides multiple methods for data producers based on application scenarios and actual data management choices. These methods include, but are not limited to, methods for forward characterizing metadata for newly added data and reverse characterizing metadata for existing data, to support proactive characterization and management of managed data. In other words, multiple modes are used to ensure compatibility with metadata characterization descriptions at different stages of system management, thereby improving the flexibility of system use.

[0056] Preferably, the preset metadata definition rules define metadata based at least on the dimensions of category, location, content, and parsing parameters. It is understood that the preset metadata definition rules include definition rules for relational metadata, file metadata, and message metadata. However, in actual application of these rules for metadata characterization, only one or a combination of these dimensions may be used to define the metadata; this is not limited here. For example, relational data may be defined based only on category, location, and content; file data may be defined based on category, location, content, and parsing parameters, with parsing parameters only used when the file data is in CSV format; message data may also be defined based only on category, location, and content. It is understood that the above examples are merely application examples of the preset metadata definition rules and do not constitute a limitation on the scope of the preset metadata definition rules.

[0057] In this embodiment, the parsing parameters are defined using a predefined approach to constrain the extended definitions of content elements. The category dimension primarily describes the business purpose of the data object, i.e., the applicable business scenario or specific data requirement; the location dimension primarily describes where the data object is obtained, defining information such as the object name; the content dimension primarily explains the specific information contained in the data object, helping to understand the business meaning of the content elements; and the parsing dimension primarily describes the method for identifying content from the target category. More preferably, for the management of content parsing parameters, this embodiment adopts a pre-configured background model. By predefining parsing parameters applicable to different formats of file-type or message-type data, the implementation cost of the parser can be effectively reduced, making the parameter parsing process more flexible.

[0058] It should be noted that the data production user in this application embodiment is preferably a person with a thorough understanding of various types of heterogeneous data. Based on the metadata definition function provided in this application embodiment, they can effectively identify and deconstruct the acquired heterogeneous data, thereby performing metadata characterization and clearly describing the business meaning and parsing parameters necessary for tailoring and processing. For example, the business semantics included in the data include, but are not limited to, connotation, attribute name, value range, attribute type, object category, and structural composition. Similarly, the technical semantics included in the data include, but are not limited to, storage location, field type, length, parsing parameters, encoding method, and expression.

[0059] Specifically, in step S1, the metadata includes newly added data and existing data;

[0060] When the metadata is newly added data, before responding to the data producer user's operation of uploading the metadata of multi-source heterogeneous data, the method further includes: receiving the data producer user's data production request, generating a data demand form and associating it with the corresponding batch number;

[0061] It's important to note that this involves metadata characterization for non-relational data in newly created applications, specifically file-based and message-based data. By receiving data production requests from users, the system parses and processes these requests according to the interface specification requirements to generate a data requirement form. A batch number is then generated online, and the requirement form and batch number are linked, enabling data traceability based on the batch number.

[0062] For example, for file-type data with JSON content format, its metadata definition interface is as follows: Figure 3 As shown. It is worth noting that, in the process of metadata characterization, to ensure a balance in the definition of object formats, this application embodiment defines format details based on general IT definitions while also providing extensibility according to specific business scenarios. For example, in defining the content of a JSON-formatted file object, in addition to adopting the existing standard data format defined by JSON, the concept of sub-elements is introduced. The sub-element type constrains the input paradigm of array data types, ensuring that data type parsing still strictly adheres to the existing JSON format definition for standardized processing, avoiding compatibility issues that may arise from introducing proprietary processing protocols.

[0063] like Figure 4 As shown, when the metadata is existing data, before the user uploads the metadata of the multi-source heterogeneous data in response to the data production user, the method further includes:

[0064] Analyze the existing data to be characterized, including basic information and content elements, to form a characterization template and publish it.

[0065] Specifically, for the management of non-relational data accessed from existing business systems, it is preferable to reduce its intrusiveness and management costs. Therefore, this embodiment adopts an active data collection and import method, that is, using the decomposition method of managed data objects during forward design to parse and format from existing data interfaces to form a characterization template table. Data production users only need to characterize the existing basic information and content elements to upload and import them in batches into the design library, thereby reducing its intrusiveness and management costs.

[0066] In step S1, the metadata is received, reviewed, and verified, and after successful verification, the metadata is converted to the runtime state, including:

[0067] The metadata is received, and the compliance and consistency of the basic information and content elements defined in the metadata are reviewed. After the review is passed, the version release verification is initiated.

[0068] The metadata within the version is archived using the version number, and the metadata is then converted to the runtime state; wherein, the version number includes at least one batch number;

[0069] When the metadata is existing data, after archiving the metadata within a version using the version number, it also includes:

[0070] Update the metadata design table and the release history table.

[0071] Specifically, when the metadata is new data, after the metadata is characterized and uploaded, the system will generate a corresponding version number based on the metadata design table, such as version V1.0. It is worth noting that there may be multiple batch numbers under the same version number, that is, there are multiple design tables under the same batch, which are combined into the same version for release. When the metadata is existing data, it is also necessary to update the original metadata design table and release the historical design table together with the current design table, such as releasing the design table of version V1.0 and the design table of version V2.0 together.

[0072] Step S2. Configure corresponding permissions for users with different roles so that each user can perform corresponding operations on the data directory based on their own permissions; among them, users with different roles include data producers, platform managers, and data consumers;

[0073] In step S2, corresponding permissions are configured for users with different roles, so that each user can perform corresponding operations on the data directory based on their own permissions, including:

[0074] Configure platform management users with at least the permissions for directory configuration management, directory classification, directory approval, and directory extension approval, so that platform management users can perform data directory planning, top-level directory creation and maintenance, data directory approval, and data directory extension approval based on their own permissions;

[0075] Configure at least the cataloging permissions for data production users so that they can expand and compile data catalogs based on their own permissions;

[0076] Configure at least directory retrieval permissions for data consumer users so that they can query data directory information based on their own permissions.

[0077] Specifically, this embodiment provides autonomous classification management of the data catalog based on the federated principle, as follows:

[0078] like Figure 5-6As shown, different roles are assigned based on production, consumption, and management scenarios, and different permissions and functions are granted to different users. The top-level design of the data directory is carried out by the directory administrator (i.e., the platform administrator), authorizing data producers to focus on their own business data. At the same time, based on the needs of data production and consumption, the directory can be expanded and managed, and data objects can be maintained. This achieves data sharing and interaction among tenants (data producers) by jointly creating a unified and standardized data directory while maintaining the independence of their business systems.

[0079] It should be noted that, preferably, the administrator and user management configuration in this embodiment has a two-level definition. Data directory management adopts a model of platform administrator + sub-directory administrator (tenant administrator) + directory producer (ordinary user), providing two-level management definitions for directory administrators and data producers. Based on different management responsibilities and data functions, combined with different usage scenarios, directory control, directory compilation, directory-related change approval, and directory query consumption are performed respectively. The specific functional division of data directory management roles is shown in the table below:

[0080]

[0081]

[0082] The directory administrator plans the data directory structure from a platform-wide perspective, creates and maintains the top-level directory, and authorizes sub-directory administrators with permissions to extend and edit directory nodes, controlling the opening of extended directories, category applications, and data directories. Tenants, as data producers, can extend subdirectories based on the data directory branches already authorized by the directory administrator after applying for data directory configuration permissions. Subdirectory extensions inherit the ownership of their parent directories. Once approved by the administrator, extended directories can be used to mount data objects or continuously expand.

[0083] Step S3. Respond to the data catalog compilation operations performed by platform management users and / or data production users, establish the association between the data catalog and data objects based on the compilation operations, and form a classified and hierarchical data catalog; wherein, the data objects are instantiated metadata objects;

[0084] It's important to note that, as the core of data catalog management, the data catalog consists of a catalog structure and data objects. The catalog structure refers to the cataloging structure information organized according to the needs of data catalog management, while data objects refer to instantiated metadata objects, containing metadata object information, corresponding data storage configurations, and related extended descriptive information. By acquiring physical information, performing logical definitions, associations, and classifications, and integrating them to form a normalized object model, it is possible to categorize and assign data to different categories as needed.

[0085] Preferably, data producers expand and compile the structure of the data directory and the mounting of data objects based on preset classification rules, classification abbreviations, and directory encoding rules.

[0086] like Figure 7 It's important to note that the directory classification information maintains global classification rules, classification abbreviations, and directory encoding rules. This provides the rule dependencies for building data framework views, organizing business data categories, and generating catalogs, serving as the primary perspective for data consumers retrieving heterogeneous data. The data directory primarily interfaces with data storage entities, organizing various data sources and classifying data entities according to business perspectives. It serves as the rule basis for data object mounting and management. Directory classification information is derived from data storage, metadata object attributes, and user-defined classification information, and is characterized by a unified directory management system. Data directory management controls the hierarchical structure and encoding rules of the data directory by specifying categories. Categories are constructed through a multi-dimensional model, forming category labels.

[0087] Preferably, in step S3, establishing the association between the data directory and the data objects based on the compilation operation includes:

[0088] Obtain the tag information of data objects under the data directory described by the data producer user, and establish the association between the data directory and the data objects under the corresponding data source based on the tag information.

[0089] It's important to note that the core of cataloging multi-source heterogeneous data is defining and characterizing data objects. This involves associating data source storage information with metadata, and using tags to refine relevant business, technical, and management attributes, forming a basic object capable of providing atomic data services to data consumers. For monolithic data objects, data producers simply need to select the corresponding metadata object, associate it with the corresponding storage type based on the metadata's structure, and specify the storage source and partitioning rules to quickly create the data object and publish it to the category directory with a single click.

[0090] Specifically, file-type metadata objects, when mounted, explicitly specify basic attributes including the server type storing the file, link address, storage directory, storage strategy, data update cycle type, expected generation time, and notification method for data generation. Message-type objects, on the other hand, explicitly specify basic attributes when mounted, including message type, data source, and message topic. When not a Kafka object, corresponding tag categories can be recorded.

[0091] In step S3, after forming the categorized and hierarchical data catalog, the method further includes:

[0092] Publish the completed data catalog and set the access method, scope, and time for the data catalog.

[0093] Preferably, after the data directory is established, this embodiment also reviews the data directory. More preferably, it provides a multi-level review and platform self-review mechanism. Data producers only need to associate metadata with storage information, enter and submit according to the agreement, and the system backend will automatically verify the integrity and logic of the configuration information based on subsequent service access needs and security policy requirements to determine whether it can be published as a complete data object. After the object is published, this embodiment will verify the hierarchical ownership permissions of the object and directory, provide open policy selection, and ensure that the open scope conforms to the permission inclusion relationship of the directory ownership level. In the process of opening or deregistering the data directory and object, operation restrictions and alarm basis are provided by providing association reference analysis and status judgment. Through the self-approval function of different stages, it is ensured that when the administrator intervenes in the review stage, he is not allowed to interfere with the logic, correctness, integrity and other issues of the configuration. It can focus on the judgment of informed consent and the introduction of external policies, which simplifies the operation of users managing the data directory.

[0094] Preferably, after the data directory structure and data objects are built and published, and after opening settings and authorization, they can enter the unified service management stage for information and data content sharing and interaction. The opening of data directories and data objects supports specifying the opening method, scope, and expiration time. The opening method and scope can be switched to globally open, restricted to non-open, or partially open to a specified tenant / role / user scope, depending on the opening status of the parent directory. Specifically, the directory administrator selects a node with a published status in configuration management, selects "open," and then obtains the opening method, opening scope, and expiration time of the nearest verified parent directory node. Details are as follows:

[0095] 1) If there is no parent node, the default initial page is globally open and never expires;

[0096] 2) If the parent node is already open = globally open, then the optional scope for initial opening is:

[0097] a) The only available opening options are partial opening and restricted opening.

[0098] b) Optional tenant / user range: All tenants / users on the platform.

[0099] c) Limit the selectable range of expiration times to less than or equal to the expiration time of the nearest parent directory node;

[0100] 3) If the parent node is already open (partially open / restricted open), then the possible range of nodes to be opened initially is:

[0101] a) The opening mode can only be selected from the corresponding options: Partial Open - Partial Open / Restricted Open - Restricted Open.

[0102] b) Optional tenant / user range: The selected range of the parent node.

[0103] c) The selectable range of expiration time restrictions is less than or equal to the expiration time of the nearest parent directory node.

[0104] Step S4. Respond to the data consumer's information retrieval operation on the data catalog and return the corresponding data attribute information and content information to the data consumer.

[0105] In step S4, responding to the data consumer's information retrieval operation on the data catalog includes:

[0106] A unified query window is used to respond to information retrieval operations by data consumers, allowing them to query information based on the data directory structure and data objects.

[0107] Specifically, this embodiment provides data consumers with access to full catalog information, detailed information viewing, and data preview, as well as a quick entry point for subscribing to target data objects. Data consumers do not need to concern themselves with the storage routing access details of objects; they can directly access the basic attributes described by metadata, such as data category, content structure, and business affiliation, from an overall perspective through the data object details. They can also view supplementary information such as data classification, physical storage, data storage type, and business tags.

[0108] Based on the above-disclosed content, this application embodiment uses preset metadata definition rules to uniformly characterize and manage multi-source heterogeneous data, applicable to business data in various formats from multiple file-type and / or message-type data sources. By providing corresponding metadata definition methods for new and existing data, data characteristics are defined to form standardized data management. Based on metadataization and data catalog abstraction and encapsulation methods, categorized and layered data assets are formed, enabling visualized data management and meeting users' needs for data cognition and retrieval. After standardizing the metadata management of file-type or message-type data, and then classifying and layering it through a data catalog, a clear data asset structure can be formed for the enterprise, allowing it to quickly and clearly understand the overall situation of its data assets and providing strong data support for subsequent data applications. For example, it can quantify and display the types, quantities, and usage status of assets, allowing enterprise managers to intuitively understand the overall situation of their assets and provide scientific data basis for capital investment and operational decisions. Through the system's ability to achieve generalized data trimming and processing, it can significantly improve the efficiency of collaborative analysis between heterogeneous data sources. Acting as a bus, it can efficiently aggregate the diverse data circulating within the enterprise and perform data alignment and statistical analysis according to demand scenarios, providing strong data support for the enterprise's business development.

[0109] The second aspect provides a unified management system for multi-source heterogeneous data based on metadata, including:

[0110] The metadata publishing module is used to respond to the operation of data producer users uploading metadata of multi-source heterogeneous data, receive, review and verify the metadata, and convert the metadata to the running state after the verification is passed; wherein, the metadata is the data obtained by data producer users based on preset metadata definition rules to characterize multi-source heterogeneous data, and the multi-source heterogeneous data includes at least relational data, file data or message data;

[0111] The permission configuration module is used to configure corresponding permissions for users with different roles, so that users with different roles can perform corresponding operations on the data directory based on their own permissions; among them, users with different roles include data producers, platform managers, and data consumers.

[0112] The data catalog publishing module is used to respond to the data catalog compilation operations of platform management users and / or data production users. Based on the compilation operations, it establishes the association between the data catalog and data objects and data source storage information, forming a classified and hierarchical data catalog; among them, data objects are instantiated metadata objects;

[0113] The data query module is used to respond to data consumer users' information retrieval operations on the data catalog and return the corresponding data attribute information and content information to the data consumer users.

[0114] Thirdly, the present invention provides a computer device comprising a memory, a processor, and a transceiver connected in sequence and communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the method described in any possible design of the first aspect.

[0115] Specifically, the memory may include, but is not limited to, Random-Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, First-In-First-Out (FIFO) Memory, and / or First-In-Last-Out (FILO) Memory, etc.; the processor may not be limited to the STM32F105 series microprocessor; the transceiver may be, but is not limited to, a WiFi (Wireless Fidelity) wireless transceiver, a Bluetooth wireless transceiver, a GPRS (General Packet Radio Service) wireless transceiver, and / or a ZigBee (a low-power LAN protocol based on the IEEE 802.15.4 standard) wireless transceiver, etc. Furthermore, the computer device may also include, but is not limited to, a power module, a display screen, and other necessary components.

[0116] The working process, working details and technical effects of the aforementioned computer device provided in the third aspect of this embodiment can be found in the method described in the first aspect or any possible design of the first aspect, and will not be repeated here.

[0117] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the method described in any possible design of the first aspect.

[0118] The computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0119] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the method described in the first aspect or any possible design of the first aspect, and will not be repeated here.

[0120] Fifthly, the present invention provides a computer program product comprising instructions that, when executed on a computer, cause the computer to perform the method described in any possible design of the first aspect.

[0121] The working process, working details and technical effects of the aforementioned computer program product containing instructions provided in the fifth aspect of this embodiment can be found in the method described in the first aspect or any possible design of the first aspect, and will not be repeated here.

[0122] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A metadata-based unified management method for multi-source heterogeneous data, characterized in that, The method comprises the following steps: In response to the operation of uploading metadata of multi-source heterogeneous data by a data production user, the metadata is received, audited and verified, and the metadata is converted to a running state after verification; wherein the metadata is data obtained by the data production user based on a preset metadata definition rule to characterize multi-source heterogeneous data, and the multi-source heterogeneous data at least includes relational data, file data or message data; Different roles of users are configured with corresponding permissions, so that each role of user can perform corresponding operations on the data directory based on their own permissions; wherein different roles of users include data production users, platform management users and data consumption users; In response to the compilation operation of the data directory by the platform management user and / or the data production user, the association between the data directory and the data object is established according to the compilation operation, and a classified and layered data directory is formed; wherein the data object is an instantiated metadata object; In response to the information retrieval operation of the data directory by the data consumption user, the corresponding data attribute information and content information are returned to the data consumption user; Different roles of users are configured with corresponding permissions, so that each role of user can perform corresponding operations on the data directory based on their own permissions, including: The platform management user is at least configured with the permissions of directory configuration management, directory classification, directory audit and directory expansion audit, so that the platform management user can plan the data directory, create and maintain the top-level directory, audit the data directory and expand the audit of the data directory based on their own permissions; The data production user is at least configured with the permission of directory compilation, so that the data production user can compile the data directory expansion based on their own permissions; The data consumption user is at least configured with the permission of directory retrieval, so that the data consumption user can query the information of the data directory based on their own permissions. 2.The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, The metadata includes new data and inventory data; When the metadata is new data, before responding to the operation of uploading metadata of multi-source heterogeneous data by a data production user, the method further comprises: receiving a data production request of the data production user, generating a data demand sheet and associating a corresponding batch number; When the metadata is inventory data, before responding to the operation of uploading metadata of multi-source heterogeneous data by a data production user, the method further comprises: Parsing the basic information and content elements to be characterized of the inventory data to form a characterization template and publish it. 3.The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, The preset metadata definition rule at least defines the metadata based on the dimensions of category, location, content and analysis parameter; wherein the definition of the analysis parameter adopts a predefinition method to constrain the expansion definition of the content element.

4. The metadata-based multi-source heterogeneous data unified management method according to claim 2, characterized in that, The metadata is received, audited and verified, and the metadata is converted to a running state after verification, including: The metadata is received, the compliance and consistency of the basic information and content elements defined in the metadata are audited, and the metadata enters the version release verification after the audit is passed; The metadata in the version is archived by using a version number, and the metadata is converted to a running state; wherein the version number at least includes a batch number; When the metadata is inventory data, after archiving the metadata in the version by using the version number, it further comprises: Update the metadata design table and the release history table.

5. The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, The data production user extends and compiles the structure of the data catalog and the mounting of the data objects based on preset classification rules, classification short codes, and catalog coding rules.

6. The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, According to the compiling operation, the association between the data catalog and the data objects is established, including: Obtain the tag information of the data objects under the data catalog described by the data production user, and establish the association between the data catalog and the data objects under the corresponding data source according to the tag information.

7. The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, After the hierarchical classification data catalog is formed, the method further includes: Release the compiled data catalog, and set the opening mode, range, and time of the data catalog. 8.The metadata-based multi-source heterogeneous data unified management method according to claim 1, characterized in that, In response to the information retrieval operation of the data consumption user on the data catalog, including: Based on the unified query window, the information retrieval operation of the data consumption user is responded to, so that the data consumption user can query information based on the data catalog structure and the data objects.

9. A metadata-based multi-source heterogeneous data unified management system, characterized in that, Including: The metadata release module is configured to respond to the operation of the data production user uploading the metadata of the multi-source heterogeneous data, receive, audit, and verify the metadata, and convert the metadata to a running state after verification; wherein the metadata is data obtained by the data production user describing the multi-source heterogeneous data based on preset metadata definition rules, and the multi-source heterogeneous data at least includes relational data, file data, or message data; The permission configuration module is configured to configure corresponding permissions for users of different roles, so that users of different roles can perform corresponding operations on the data catalog based on their own permissions; wherein the users of different roles include data production users, platform management users, and data consumption users; Configuring corresponding permissions for users of different roles so that users of different roles can perform corresponding operations on the data catalog based on their own permissions, including: The platform management user is configured with at least the permissions of catalog configuration management, catalog classification, catalog audit, and catalog extension audit, so that the platform management user can perform data catalog planning, top-level catalog creation and maintenance, data catalog audit, and data catalog extension audit based on his own permissions; The data production user is configured with at least the permission of catalog compilation, so that the data production user can perform data catalog extension compilation based on his own permissions; The data consumption user is configured with at least the permission of catalog retrieval, so that the data consumption user can query the data catalog information based on his own permissions; The data catalog release module is configured to respond to the compiling operation of the platform management user and / or the data production user, establish the association between the data catalog and the data object and the data source storage information according to the compiling operation, and form a hierarchical classification data catalog; wherein the data object is an instantiated metadata object; The data query module is configured to respond to the information retrieval operation of the data consumption user on the data catalog, and return the corresponding data attribute information and content information to the data consumption user.

Citation Information

Patent Citations

  • System and method for realizing data sharing service based on multi-view data directory

    CN112783931A

  • Resource cataloguing method and device

    CN113342921A