A data processing method and apparatus

By defining a unified metadata format, the data integration problem caused by the independence of directory services between data lakes is solved, and simple communication and efficient data transmission between data systems are achieved.

CN114902636BActive Publication Date: 2025-06-13HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980103358.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-31
Publication Date
2025-06-13
Estimated Expiration
2039-12-31

AI Technical Summary

Technical Problem

The directory services between existing data lakes are independent of each other, resulting in different data lake products of different manufacturers having different directory services, which cannot be directly integrated, and the metadata format conversion is complex.

Method used

Simplify communication between two data systems by defining the metadata format that both data systems can recognize. The specific method includes obtaining the metadata of the data in the first data system, converting it into a unified format, and then converting it into the metadata format of the second data system after transmission.

Benefits of technology

It realizes data transmission between different data systems, simplifies the communication process, reduces the complexity of metadata format conversion, and improves the efficiency of data integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902636B_ABST
    Figure CN114902636B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a data processing method and apparatus. The method includes: obtaining metadata of data in a first data system; converting the metadata of the data in the first data system from a first format to a second format, where the first format is the format of the metadata of the data in the first data system, and the second format is a metadata format that can be recognized by both the first data system and a second data system. By using the method provided by the embodiment of the present invention, a unified metadata format that can be recognized by multiple data systems is defined for multiple data systems, which simplifies the data transmission between different data systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the storage field, and in particular, to a data processing method and apparatus. Background Art

[0002] A data lake is a repository or system that stores data in its original format without prior structuring of the data. A data lake can store structured data (such as tables in a relational database), semi-structured data (such as CSV, logs, XML, JSON), unstructured data (such as emails, documents, PDFs), and binary data (such as graphics, audio, video). To manage the data in a data lake, it is required that all data entering the data lake must provide relevant metadata, and the metadata of different types of data is organized into a directory service. Subsequently, when using the data, the data can be analyzed based on the directory service to provide the analyzed data to users.

[0003] Currently, the directory services in each data lake are independent of each other, and the directory services adopted by each data lake come from different manufacturers. For example, in the security field, each province or city purchases data lake products from different manufacturers, and the data lake products of different manufacturers have different directory services. Therefore, when a higher-level unit, such as a province or the central government, wants to integrate the data of a lower-level unit, such as a city or a province, it cannot directly integrate because it cannot recognize the directory services of each city or province. One current implementation method is as follows: If the directory services of data lake A (such as a province) and data lake B (such as a city) are different, but data lake A needs to obtain the data of data lake B, then data lake A needs to notify data lake B and tell data lake B the data it wants to access. Data lake B places the metadata of the data that data lake A wants to access on an exchange platform. The exchange platform converts the metadata of the data that data lake A wants to access into a directory service that data lake A can recognize, and notifies data lake A to obtain the metadata from the exchange platform, and then obtains the data in data lake B based on the metadata.

[0004] In the above implementation method, it is necessary to directly convert the format of the metadata defined in the directory service of data lake A into the format of the metadata defined in the directory service of data lake B. In this way, when data lake B needs to obtain the data of multiple data lakes with different directory services, the exchange platform needs to know the directory services of all data lakes, resulting in a relatively high complexity of metadata format conversion. Summary of the Invention

[0005] The present invention provides a data processing method and apparatus. The present invention simplifies the communication between two data systems by defining a metadata format that can be recognized by both data systems.

[0006] In a first aspect of the present invention, a data processing method is provided. The method includes: obtaining metadata of data in a first data system; converting the metadata of the data in the first data system from a first format to a second format, where the first format is the format of the metadata of the data in the first data system, and the second format is a metadata format recognizable by both the first data system and a second data system.

[0007] In the data processing method provided by the embodiments of the present invention, by defining a unified metadata format for different data systems, when the first data system needs to obtain data from the second data system, the second data system first converts the metadata format of the second data system to the metadata format defined by the directory template. After the first data system obtains the metadata converted to the unified format, it then converts it to the format of the metadata in the first data system. In this way, as long as each data system can recognize the unified metadata format, it can perform data transmission with other data lakes, thus simplifying communication between different data systems.

[0008] In a possible implementation of the first aspect, the method further includes pre - defining a metadata template, and the metadata format adopted by the metadata template is the second format.

[0009] In a possible implementation manner of the first aspect, the method further includes: sending the metadata converted to the second format to a message platform so that the second data system can obtain the metadata converted to the second format from the message platform.

[0010] The second data system sends the metadata to the message platform. In this way, the first data system can obtain the metadata of the second system from the message platform, and thus does not need to actively obtain the metadata from the second data system.

[0011] In a possible implementation manner of the first aspect, in the step of sending the metadata converted to the second format to the message platform, it includes: determining whether the metadata converted to the second format meets a preset publishing rule;

[0012] When it is determined that the metadata converted to the second format meets the preset publishing rule, the metadata converted to the second format is sent to the message platform.

[0013] By setting the publishing rule of the metadata in the second data system, only the metadata that meets the publishing rule will be published to the message platform. In this way, without the need for permission authentication, the security of the data can also be ensured.

[0014] In a possible implementation manner, the method further includes:

[0015] Send a request to create a publishing metadata queue to the message platform so that the message platform establishes a publishing metadata queue;

[0016] Send the metadata converted to the second format to the message platform, including:

[0017] Write the metadata converted to the second format into the publishing metadata queue of the message platform.

[0018] A data federation is formed between the first data system and the second data system, and a publishing metadata queue is established. A data transmission channel between the first data system and the second data system is established through the publishing metadata queue, optimizing the data transmission process.

[0019] In a possible implementation, the method further includes:

[0020] Obtain the address of the publishing metadata queue sent by the message platform, and send the address of the publishing metadata queue to the second data system.

[0021] In a possible implementation, the method further includes:

[0022] Send a request to create an original metadata queue to the message platform so that the message platform establishes an original metadata queue;

[0023] Converting the format of the metadata in the first data system from the first format to the second format includes:

[0024] Convert the metadata of the first data system in the original metadata queue from the first format to the second format.

[0025] By establishing an original metadata queue, caching can be provided for the converted data, ensuring data reliability.

[0026] A second aspect of the present invention provides a data processing method, the method includes: obtaining the metadata of the data of the first data system; converting the format of the data metadata of the first data system from the second format to the third format, where the second format is a metadata format recognizable by both the first data system and the second data system, and the third format is the metadata format of the data in the second data system.

[0027] In the data processing method provided by the embodiments of the present invention, by defining a unified metadata format for different data systems, when the first data system needs to obtain data from the second data system, the second data system first converts the metadata format of the second data system into the metadata format defined by the directory template. After the first data system obtains the metadata converted into the unified format, it then converts it into the metadata format in the first data system. In this way, as long as each data system can recognize the unified metadata format, it can perform data transmission with other data lakes, thus simplifying the communication between different data systems.

[0028] In a possible implementation manner, the method includes pre-defining a metadata template, and the metadata format adopted by the metadata template is the second format.

[0029] In a possible implementation manner, the obtaining of the metadata of the first data system includes:

[0030] Obtaining the metadata of the data of the first data system from the message platform.

[0031] By obtaining the metadata of the first data system through the message platform, the second data system can obtain the metadata of the first system without accessing the first system.

[0032] In a possible implementation manner, the method further includes:

[0033] Judging whether the metadata meets a preset receiving rule;

[0034] When it is determined that the converted metadata meets the preset receiving rule, converting the format of the metadata of the data of the first data system from the first format to the second format.

[0035] By setting the receiving rule, data that the second data system does not need can be filtered out, thus ensuring data security.

[0036] In a possible implementation manner, the method further includes:

[0037] Receiving the address of the published metadata queue in the message platform sent by the first data system;

[0038] The obtaining of the metadata of the data of the first data system from the message platform includes:

[0039] Obtaining the metadata from the published metadata address according to the address of the published metadata queue.

[0040] The third aspect of the present invention provides a data processing device, which includes a plurality of modules for performing the functions of the steps in the data processing method of the first aspect. The beneficial effects achieved are the same as those of the methods provided in the first aspect and will not be elaborated here.

[0041] The fourth aspect of the present invention provides a data processing device, which includes a plurality of modules for performing the functions of the steps in the data processing method of the second aspect. The beneficial effects achieved are the same as those of the methods provided in the first aspect and will not be elaborated here.

[0042] The fifth aspect of the present invention provides a server, which includes a processor and a memory. Program instructions are stored in the memory, and the processor executes the program instructions to implement the functions of the method provided in the first aspect.

[0043] The sixth aspect of the present invention provides a server, which includes a processor and a memory. Program instructions are stored in the memory, and the processor executes the program instructions to implement the functions of the method provided in the second aspect.

[0044] The seventh aspect of the present invention provides a storage medium, in which program instructions are stored, and the program instructions are executed by a processor to implement the functions of the method provided in the first aspect.

[0045] The eighth aspect of the present invention provides a storage medium, in which program instructions are stored, and the program instructions are executed by a processor to implement the functions of the method provided in the second aspect. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0047] Figure 1 It is a schematic diagram of the structure of the data lake provided by the embodiment of the present invention.

[0048] Figure 2 It is the format of the metadata in the directory template defined in the embodiment of the present invention.

[0049] Figure 3 It is the structure diagram of the federated system in the embodiment of the present invention.

[0050] Figure 4 It is the flowchart of the method for registering data lakes as a federation in the embodiment of the present invention.

[0051] Figure 5Flow chart of the method for data transmission between data lakes in the embodiments of the present invention.

[0052] Figure 6 Metadata format of the Hive database in the embodiments of the present invention.

[0053] Figure 7 Schematic diagram showing that the metadata format of the Hive database in one data lake is converted into the metadata format defined by the directory template in the embodiments of the present invention.

[0054] Figure 8 Schematic diagram showing that the metadata format defined by the directory template is converted into the metadata format in another data lake in the embodiments of the present invention.

[0055] Figure 9 Schematic diagram of the application scenario of forming a data federation for multiple data lakes in the embodiments of the present invention.

[0056] Figure 10 Functional module diagram of the first data processing device and the second data processing device in the embodiments of the present invention.

[0057] Figure 11 Structure diagram of the server on which the federation service runs in the embodiments of the present invention. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0059] Metadata: Also known as intermediary data or relay data, it is data that describes data, mainly information describing data attributes, and is used to support functions such as indicating storage locations, historical data, resource searches, file records, etc. Metadata is a kind of electronic catalog. In order to achieve the purpose of compiling a catalog, it is necessary to describe the attributes of the data, and thus achieve the purpose of assisting data retrieval.

[0060] Metadata format: Defines the attributes included in the metadata, the expression methods of each attribute, and the arrangement order between the attributes.

[0061] The data lake provided in the embodiments of the present invention can be deployed in a server cluster or in a cloud environment. As Figure 1 shown, it is a schematic diagram of the structure of the data lake deployed in the cloud environment.

[0062] The data lake mainly includes two parts. One part is the computing resource 10, and the other part is the storage resource 20. The storage resource 20 generally includes multiple storage devices 201, and the data stored in the data lake is generally distributed and stored in multiple storage devices 201. The computing resource 10 generally uses virtual machines 101 in a cloud environment, and the virtual machines 101 are deployed in computing servers (not shown in the figure). When data is stored in the data lake, the virtual machine 101 determines the storage device where the data is to be stored, obtains the metadata of the data, and records the metadata in the directory service 102.

[0063] If the data lake is deployed in a server cluster, the computing resource of the data lake is at least one computing server, and the storage resource is a storage server.

[0064] Currently, many manufacturers in the market provide data lake products, and the directory services set by different manufacturers' data lake products are not the same. However, in some scenarios, it is necessary for the data between data lakes to flow to each other. For example, when a higher-level unit such as a province or the central government wants to integrate the data of a lower-level unit such as a city or a province, since the directory services of each city or province cannot be recognized, direct integration cannot be carried out. One current implementation method is as follows: If the directory services of data lake A (such as a province) and data lake B (such as a city) are different, but data lake A needs to obtain the data of data lake B, then data lake A needs to notify data lake B and tell data lake B the data it wants to access. Data lake B needs to confirm whether data lake A has the permission to access the data it wants to access. If so, it will put the data and its metadata that data lake A wants to access on the exchange platform. The exchange platform converts the metadata of the data that data lake A wants to access into metadata that data lake A can recognize, and notifies data lake A to obtain the metadata from the exchange platform, and based on the obtained metadata, obtains the data from data lake A.

[0065] In the above implementation method, it is necessary to directly convert the format of the metadata defined in the directory service of data lake A into the format of the metadata defined in the directory service of data lake B. In this way, when data lake B needs to obtain the data of multiple data lakes with different directory services, the exchange platform needs to know the directory services of all data lakes, resulting in a relatively high complexity of metadata format conversion.

[0066] In addition, data lake A needs to actively obtain the data of data lake B, so data lake A needs to know in advance what data is in data lake B. When the scale of the data lake is relatively large, it is very difficult to predict what data is in data lake B.

[0067] In addition, every time data lake A obtains the data of data lake B, data lake B also needs to perform permission authentication on data lake A, thus affecting the data access efficiency.

[0068] In the data processing method provided by the embodiment of the present invention, a directory template is set. This directory template defines a unified metadata format for data lakes using different directory services, and all data lakes can recognize the metadata format defined in the directory template. In this way, when data lake A needs to obtain the data of data lake B, data lake B first converts the metadata format in the directory service of data lake B into the metadata format defined by the directory template. After data lake A obtains the metadata converted into the format defined by the directory template, it then converts it into the metadata format in the directory service of data lake A. In this way, as long as each data lake can recognize the metadata format defined in the directory template, it can perform data transmission with other data lakes, thus simplifying the communication between different data lakes.

[0069] In addition, data lake A and data lake B will form a data federation. After forming the data federation, when data lake B generates new metadata, data lake B will publish the newly generated metadata to the message platform. After the message platform receives the metadata published by data lake B, it will push the received metadata to data lake A. In this way, data lake A does not need to actively obtain metadata from data lake B.

[0070] In addition, a metadata publishing rule is set in data lake B, and only the metadata that meets the publishing rule will be published to the message platform. A metadata subscription rule is also set in data lake A, and only the metadata that meets the subscription rule will be stored in data lake A. In this way, data security can be ensured without the need for permission authentication.

[0071] Since the data lake includes various types of data, such as the Hive database and Hbase database belonging to structured data, and videos, audios, etc. belonging to unstructured data, different metadata formats are set in the directory service according to different types of data. For different manufacturers, when establishing the directory service of the data lake, the description formats of the metadata for each type of data will be different. Therefore, the embodiment of the present invention provides a directory template for converting the metadata formats in the directory services of different manufacturers into a unified metadata format.

[0072] As Figure 2 shown, it is a schematic diagram of the metadata format of the Hive database in the directory template defined in the embodiment of the present invention.

[0073] In Figure 2In the template shown, the expression method and order are defined for each attribute of the metadata. For example, if the table name of the Hive metadata in the directory service of data lake A is represented by TalX and the column is represented by ColX, and the table name of the Hive metadata in the directory service of data lake B is represented by TalY and the column is represented by ColX, then after being converted into the format defined by the directory template, the table name is represented by Name and the column is represented by Column.

[0074] As Figure 3 shown, it is a schematic diagram of two data lakes forming a federation. In the embodiment of the present invention, the data lake 30 and the data lake 31 simultaneously provide a directory service 301 and a federation service 302. The directory service 301 and the federation service 302 can be provided by different virtual machines in the cloud environment, or can be different processes in the same virtual machine; in the server cluster, they can be provided by different servers, or can be different processes in the same server. The message platform 33 is used to provide a queue service for the data lake 30 and the data lake 31, and provide a raw metadata queue 331 and a published metadata queue 332 for the data lakes 30 and 31, so as to facilitate the metadata exchange between the data lakes 30 and 31. The message platform can be a virtual machine cluster composed of multiple virtual machines, or a server cluster composed of multiple servers.

[0075] Please also refer to Figure 4 , which is a flowchart of the method for the data lake 30 to register with the data lake 31 when it needs to form a data federation with the data lake 31 to obtain the data in the data lake 31.

[0076] Step S401, the federation service 302 of the data lake 30 sends a registration request to the virtual machine (which can also be a server or a process providing the federation service, hereinafter referred to as the federation service) of the data lake 31 that provides the federation service.

[0077] In the embodiment of the present invention, the application program of the federation service will be installed in the data lake. When the data lake 30 needs to obtain the data in the data lake 31, the user will start the application program of the federation service 302, select the data lake 31 with which the federation needs to be established, and then send a registration request to the data lake 31 through the registration function provided by the federation service. The address of the federation service of the data lake 30 is carried in the registration request.

[0078] Step S402, when the federation service 312 of the data lake 31 receives the registration request, it respectively sends a raw metadata queue creation request and a published metadata queue creation request to the message platform 33, and at the same time sends the addresses of the federation service 312 and the federation service 302 to the message platform 33.

[0079] Step S403: The message platform 33 receives the original metadata queue creation request and the published metadata queue creation request, and creates an original metadata queue and a published metadata queue respectively.

[0080] The original metadata queue and the published metadata queue are processes running in the message platform 33. When the message platform creates the original metadata queue and the published metadata queue, a certain amount of storage space is allocated for each queue. The original metadata queue is used to store the newly generated metadata in the data lake 30, and the published metadata queue is used to store the metadata that the data lake 30 needs to publish. The process of storing metadata in the original metadata queue and the published metadata queue will be described below.

[0081] Step S404: After the original metadata queue and the published metadata queue are created, the message platform 33 returns the addresses of the original metadata queue and the published metadata queue to the federated service 312. The federated service 312 sends the address of the published metadata queue to the federated service 302 and returns the address of the original metadata queue to the directory service 311. In this way, the federated service 302 or the directory service 311 can obtain metadata from the published metadata queue or the original metadata queue respectively according to the address. In this way, a data federation is formed between the data lake 30 and the data lake 31.

[0082] The following will combine Figure 5 the process of transmitting data between the data lakes that form the federation.

[0083] As Figure 5 shown, after the data lake 30 and the data lake 31 form a data federation, metadata can be transmitted between the data lake 30 and the data lake 31. The process of data transmission is as Figure 5 shown.

[0084] Step S501: The virtual machine providing the directory service 311 (which can also be a process or a server, hereinafter all referred to as the directory service) detects the newly generated metadata in the directory service in the data lake 31.

[0085] When the user generates a new table or file in the data lake 31, the directory service 311 obtains the metadata of the newly generated table or file and stores it. In the embodiment of the present invention, a plugin is added to the directory service to detect whether the directory service has obtained new metadata.

[0086] Step S502: The directory service 311 sends a metadata write request to the message platform 33 according to the address of the original metadata queue for the newly generated metadata.

[0087] When establishing a data federation, the federation service 312 sends the address of the original metadata queue to the directory service. Therefore, when the directory service 311 generates new metadata, a metadata write request can be generated to send the metadata write request to the message platform 33.

[0088] Step S503, the message platform 33 writes the newly generated metadata into the original metadata queue.

[0089] Step S504, the message platform 33 notifies the federation service 312 of the data lake 31 that there is new source data generated in the original metadata queue.

[0090] Step S505, the federation service 312 obtains the metadata from the original metadata queue, and then converts the format of the metadata into the format of the metadata defined in the directory template.

[0091] For example, if the metadata is the metadata in a Hive database, in the data lake 31, the expression of its metadata is as Figure 6 shown. After converting it to the format of the Hive metadata defined in the directory template as Figure 2 shown, it is as Figure 7 shown. The corresponding relationship between each attribute in the metadata template and each attribute in the Hive metadata is predefined in the federation service 312. When converting, first read the Figure 2 shown metadata template, then obtain the value corresponding to the attribute that is the same as the attribute in the metadata template in the Hive metadata, and fill the obtained value into the corresponding position of the metadata template. For example, if the attribute "Name" in the Hive metadata represents the same attribute as the attribute "Table Name" in the metadata template, then obtain the value "Hive-table1" of the attribute "Name" from the Hive metadata and write it into the attribute "Table Name" defined in the metadata template. Similar operations are performed for other attributes, and the Figure 6 shown metadata format of the Hive database in the data lake 31 can be converted to the Figure 7 shown format defined by the metadata template. In addition, the values of some attributes can also be simplified, and only the key information needs to be extracted. For example, for the Figure 6 DB attribute, the expression in the data lake 31 is relatively complex and also includes the global identifier of the metadata identifying the database to which this table belongs. Then, when converting to the format of the metadata template, only extract the database identifier "hive_db".

[0092] Step S506, determine whether the metadata can be published to the federated data lake 30 of the data lake 31 according to the publishing rule.

[0093] In the embodiments of the present invention, the metadata published to the federated data lake 30 is filtered, and only the data that meets the publishing rules will be published to the federated data lake. The publishing rules are set by the user according to the attributes of the metadata in the catalog template. The publishing rules include rules and actions.

[0094] For example, the rule of the publishing rule is set as: creator = Zhang San & generation time > 2019 / 9 / 18 & table type = finance, and the action is set as publish, indicating that the data that meets the rule can be published.

[0095] In the rule, the creator, generation time, and table type are all attributes defined in the metadata template. According to the rule, after Zhang San created a finance-related table after September 18, 2019, the metadata information related to this table will be synchronized to other data lakes that form the federation.

[0096] Another example, the rule of the publishing rule is set as: metadata source = XX Company, and the action is set as not publish, then if the metadata is from XX Company, it will not be synchronized to the federated data lake 30 of the data lake 31.

[0097] Step S507, when it is determined according to the publishing rule that the metadata can be published to the federated data lake 30 of the data lake 31, the federated service 312 sends the metadata to the message platform 33.

[0098] When registering as the federated data lake of the data lake 31 in the data lake 30, the data lake 31 applies to establish a publishing metadata queue in the message platform 33 and returns the address of the publishing metadata queue to the federated service 312 of the data lake 31. In this way, when the federated service 312 sends the metadata to the message platform 33, it can carry the address of the publishing metadata queue.

[0099] Step S508, the message platform 33 stores the received metadata in the publishing metadata queue.

[0100] Step S509, the message platform 33 notifies the federated service 302 of the data lake 30 to obtain data from the publishing metadata queue.

[0101] During the registration process of the federated data lake, the federated service 312 passes the address of the federated service 302 to the message platform. Therefore, the message platform can notify the federated service 302 to obtain the metadata according to the address of the federated service 302.

[0102] Step S510, the federated service 302 of the data lake 30 obtains the metadata from the publishing metadata queue according to the address of the publishing metadata queue.

[0103] During the registration process of the federated data lake, the federated service 312 sent the address of the published metadata queue to the federated service 302, so that the federated service 302 can obtain the metadata from the published metadata queue according to the address of the published metadata queue.

[0104] Step S511, the federated service 302 determines whether the metadata is the data required by the data lake B according to the subscription rules.

[0105] The subscription rules are also set according to the attributes of the metadata in the catalog template. The subscription rules also include two parts: rules and actions. If the obtained metadata meets the set rules, the corresponding actions are executed.

[0106] For example, the rule of the subscription rule can be set as: metadata source = XX, and the action is set to not receive, which means that if the metadata obtained from the data lake A is XX, the metadata will not be received.

[0107] By setting the subscription rules, the data that the data lake 30 does not need can be filtered out.

[0108] Step S512, the federated service 302 converts the metadata into the format of the metadata in the data lake 30.

[0109] In the federated service 302, the corresponding relationship between each attribute of the metadata in the data lake 30 and each attribute of the metadata in the metadata template is defined. Through the defined corresponding relationship, the metadata format in the metadata template can be converted into the metadata format in the data lake 30. After conversion, as Figure 8 shown.

[0110] Step S513, the federated service 302 sends the converted metadata to the catalog service 301 of the data lake 30.

[0111] Step S514, the catalog service 301 of the data lake 30 stores the converted metadata in the catalog service of the data lake 30.

[0112] In the data processing method provided by the embodiment of the present invention, a directory template is set. This directory template defines a unified metadata format for data lakes using different directory services, and all data lakes can recognize the metadata format defined in the directory template. In this way, when data lake A needs to obtain the data of data lake B, data lake B first converts the metadata format in the directory service of data lake B into the metadata format defined by the directory template. After data lake A obtains the metadata converted into the format defined by the directory template, it then converts it into the metadata format in the directory service of data lake A. In this way, as long as each data lake can recognize the metadata format defined in the directory template, it can perform data transmission with other data lakes, thus simplifying the communication between different data lakes.

[0113] In addition, data lake A and data lake B will form a data federation. After forming the data federation, when data lake B generates new metadata, data lake B will publish the newly generated metadata to the message platform. After the message platform receives the metadata published by data lake B, it will push the received metadata to data lake A. In this way, data lake A does not need to actively obtain metadata from data lake B.

[0114] In addition, a metadata publishing rule will be set in data lake B, and only the metadata that meets the publishing rule will be published to the message platform. A metadata subscription rule is also set in data lake A, and only the metadata that meets the subscription rule will be stored in data lake A. In this way, without the need for permission authentication, the security of the data can also be ensured.

[0115] Since the data lake includes various types of data, such as the Hive database and Hbase database belonging to structured data, and videos, audios, etc. belonging to unstructured data, different metadata formats are set in the directory service according to different types of data. For different manufacturers, when establishing the directory service of the data lake, the description formats of the metadata for each type of data will be different. Therefore, the embodiment of the present invention provides a directory template for converting the metadata formats in the directory services of different manufacturers into a unified metadata format.

[0116] In practical applications, data lake 31 can be a data lake of a lower-level unit such as a city or a province, while data lake 30 is a data lake of a higher-level unit such as a province or the central government. In this way, since data lake 30 can directly obtain data from data lake 31, it is very convenient to integrate the data of lower-level units.

[0117] Such as Figure 9As shown in the figure, it is an application scenario of an embodiment of the present invention. The data lake B forms a directory service federation with the data lake D and the data lake E respectively. Then, the data lake B can obtain the metadata in the data lake D and the data lake E respectively. The data lake A forms a directory service federation with the data lake B and the data lake C respectively. Then, the data lake A can obtain the metadata in the data lake D and the data lake E respectively. Suppose the data lake D and the data lake E are data lakes of municipal units, the data lake B and the data lake C are data lakes of provincial units, and the data lake A is a data lake of a central unit. Then, the superior unit can conveniently integrate the metadata of the subordinate units, and the directory services between units at the same level are isolated from each other.

[0118] The above embodiments are only described by taking the data lake as an example, but the present invention is equally applicable to other data systems, such as databases, data warehouses, etc. Databases and data warehouses will have metadata, so the implementation methods are basically the same as those of the data lake, and will not be elaborated here.

[0119] As Figure 10 shown in the figure, it is a functional module diagram of the first data processing device 1001 and the second data processing device 2001. The first data processing device 1001 and the second data processing device 2001 are mainly used to implement the functions of the federation service 302 and the federation service 312.

[0120] The first data processing device 1001 includes a definition module 1002, a creation module 1003, an acquisition module 1004, a conversion module 1005, and a publishing module 1006. The second data processing device includes a definition module 2001, an acquisition module 2002, and a conversion module 2003.

[0121] The definition module 1002 and the definition module 2002 are respectively used to define the format of the metadata in the directory template shown in Figure 2 . For the specific definition method, please refer to the description of Figure 2 .

[0122] The creation module 1003 is used to create an original metadata queue and a published metadata queue in the message platform after receiving a registration request from the second data processing device 2001. For the specific creation process, please refer to the relevant description of Figure 4 , and will not be elaborated here.

[0123] The acquisition module 1004 is used to acquire the metadata of the data in the data lake 30. For the specific acquisition method, please refer to the description of steps S501 to S505 in Figure 5 , and will not be elaborated here.

[0124] The conversion module 1005 is used to convert the format of the metadata of the data in the data lake 30 into the metadata format defined in the catalog template. For the specific conversion method, please refer to Figure 5 the description of step S505 in

[0125] The publishing template 1006 is used to publish the metadata that has been converted into the metadata format defined in the catalog template to the publishing metadata queue of the message platform. For details, please refer to Figure 5 the descriptions of steps 507 and 508 in

[0126] After receiving the notification from the message platform, the acquisition module 2003 of the second data processing device 2001 acquires the metadata that has been converted into the format defined by the catalog template from the publishing metadata queue of the message platform. For the specific conversion process, please refer to the relevant description of step S510, which will not be elaborated here.

[0127] The conversion module 2004 is used to filter the acquired metadata according to the preset filtering rules. When it is determined that the metadata is data that the data lake 31 can receive, the format of the acquired metadata is converted from the format defined by the catalog template into the metadata format of the data in the data lake 31, and the converted metadata is stored in the data catalog of the data lake 31. For details, please refer to the descriptions of steps S511 to S514, which will not be elaborated here.

[0128] When any of the above modules or units is implemented in software, the software exists in the form of computer program instructions and is stored in the memory. The processor can be used to execute the program instructions to implement the above method flow. The processor may include, but is not limited to, at least one of the following: central processing unit (CPU), microprocessor, digital signal processor (DSP), microcontroller unit (MCU), or various computing devices for running software such as artificial intelligence processors. Each computing device may include one or more cores for executing software instructions to perform operations or processing. The processor can be a single semiconductor chip or integrated with other circuits into a semiconductor chip. For example, it can form a SoC (system on a chip) with other circuits (such as codec circuits, hardware acceleration circuits, or various bus and interface circuits), or can be integrated as an embedded processor of an ASIC in the ASIC. The ASIC integrated with the processor can be separately packaged or packaged together with other circuits. In addition to the cores for executing software instructions to perform operations or processing, the processor may further include necessary hardware accelerators, such as field programmable gate array (FPGA), PLD (programmable logic device), or logic circuits for implementing dedicated logical operations.

[0129] When the above modules are implemented in hardware circuits, the hardware circuits may be implemented by a general-purpose CPU (Central processing unit, central processor), MCU (Micro controller Unit, microcontroller), MPU (Microprocessing unit, microprocessor), DSP (Digital signal processing, digital signal processor), SoC (System on Chip, system on a chip), and of course, can also be implemented by an application-specific integrated circuit (ASIC), or a programmable logic device (programmable logic device, PLD). The above PLD can be a complex programmable logic device (complex programmable logical device, CPLD), field-programmable gate array (field-programmable gate array, FPGA), generic array logic (generic array logic, GAL), or any combination thereof, which can run the necessary software or execute the above method flow without relying on software.

[0130] As shown Figure 11 in the figure, it is a hardware structure diagram of a server. The server is a server for running the federated service in the data lake 30 or the data lake 31. The server 1101 includes a processing unit 1102, a storage unit 1103, and a communication unit 1104.

[0131] The processing unit 1101 has various implementation forms. For example, it can be a central processing unit (CPU) or a graphics processing unit (GPU), and can be a single-core processor or a multi-core processor. If implemented as a multi-core processor, the number of processors is not limited.

[0132] The storage unit 1103 can be a Dynamic Random Access Memory (DRAM) or a storage class memory (SCM). The storage unit 1103 is used to store program codes and data so that the processing unit can call and execute them to implement related functions. In the embodiments of the present invention, the program codes corresponding to the federated service can be stored, and the processor calls the program codes to execute Figure 5 the functions performed by the federated service of the data lake 30 or the functions performed by the federated service of the data lake 31 in the method shown.

[0133] The communication unit 1104 is used to communicate with the message platform and other servers for data transmission.

[0134] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, the method includes: obtaining metadata of data in a first data system; converting the metadata of data in the first data system from a first format to a second format, where the first format is the format of the metadata of data in the first data system, and the second format is a metadata format recognizable by both the first data system and a second data system; sending the metadata converted to the second format to a message platform, so that the second data system can obtain the metadata converted to the second format from the message platform.

2. The data processing method according to claim 1, characterized in that, it further includes: predefining a metadata template, and the metadata format adopted by the metadata template is the second format.

3. The data processing method according to claim 1, characterized in that, the sending the metadata converted to the second format to the message platform includes: judging whether the metadata converted to the second format meets a preset publishing rule; when it is determined that the metadata converted to the second format meets the preset publishing rule, sending the metadata converted to the second format to the message platform.

4. The data processing method according to claim 1, characterized in that, it further includes: sending a publishing metadata queue creation request to the message platform, so that the message platform establishes a publishing metadata queue; sending the metadata converted to the second format to the message platform includes: writing the metadata converted to the second format into the publishing metadata queue of the message platform.

5. The data processing method according to claim 4, characterized in that, it further includes: obtaining the address of the publishing metadata queue sent by the message platform, and sending the address of the publishing metadata queue to the second data system.

6. The data processing method according to claim 4, characterized in that, it further includes: sending an original metadata queue creation request to the message platform, so that the message platform establishes an original metadata queue; converting the format of the metadata in the first data system from the first format to the second format includes: converting the metadata of the first data system in the original metadata queue from the first format to the second format.

7. A data processing method, characterized in that, the method includes: obtaining the metadata of the data of the first data system from the message platform; converting the format of the data metadata of the first data system from the second format to a third format, where the second format is a metadata format recognizable by both the first data system and the second data system, and the third format is the format of the metadata of the data in the second data system.

8. The data processing method according to claim 7, characterized in that, it further includes: predefining a metadata template, and the metadata format adopted by the metadata template is the second format.

9. The data processing method according to claim 7, characterized in that, it further includes: judging whether the metadata meets a preset receiving rule; After determining that the transformed metadata meets the preset reception rules, convert the format of the metadata of the data in the first data system from the second format to the third format.

10. The data processing method according to claim 7, characterized in that, further comprising: receiving the address of the published metadata queue in the message platform sent by the first data system; The obtaining the metadata of the data of the first data system from the message platform includes: obtaining the metadata from the published metadata address according to the address of the published metadata queue.

11. A data processing apparatus, characterized in that, the apparatus includes: an obtaining module, configured to obtain metadata of data in a first data system; a conversion module, configured to convert the metadata of the data in the first data system from a first format to a second format, where the first format is the format of the metadata of the data in the first data system, and the second format is a metadata format recognizable by both the first data system and a second data system; a publishing module, configured to send the metadata converted to the second format to the message platform, so that the second data system obtains the metadata converted to the second format from the message platform.

12. The data processing apparatus according to claim 11, characterized in that, further comprising: a definition module, configured to pre-define a metadata template, and the metadata format adopted by the metadata template is the second format.

13. The data processing apparatus according to claim 11, characterized in that, the publishing module is further configured to: judge whether the metadata converted to the second format meets a preset publishing rule; when it is determined that the metadata converted to the second format meets the preset publishing rule, send the metadata converted to the second format to the message platform.

14. The data processing apparatus according to claim 11, characterized in that, further comprising: a creation module, configured to send a published metadata queue creation request to the message platform, so that the message platform establishes a published metadata queue; The publishing module is specifically configured to: write the metadata converted to the second format into the published metadata queue of the message platform.

15. The data processing apparatus according to claim 14, characterized in that, the creation module is specifically configured to: obtain the address of the published metadata queue sent by the message platform, and send the address of the published metadata queue to the second data system.

16. The data processing apparatus according to claim 14, characterized in that, the creation module is further configured to: send an original metadata queue creation request to the message platform, so that the message platform establishes an original metadata queue; The conversion module is specifically configured to: convert the metadata of the first data system in the original metadata queue from the first format to the second format.

17. A data processing apparatus, characterized in that, the apparatus includes: an obtaining module, configured to obtain metadata of data of a first data system from a message platform; A conversion module, configured to convert the format of the data metadata of the first data system from a second format to a third format, where the second format is a metadata format recognizable by both the first data system and the second data system, and the third format is the metadata format of the data in the second data system.

18. The data processing device according to claim 17, wherein, it further includes: A definition module, configured to pre-define a metadata template, and the metadata format adopted by the metadata template is the second format.

19. The data processing device according to claim 17, wherein, the conversion module is further configured to: Determine whether the metadata meets a preset reception rule; When it is determined that the converted metadata meets the preset reception rule, convert the format of the metadata of the data of the first data system from the second format to the third format.

20. The data processing device according to claim 17, wherein, the obtaining module is specifically configured to: Receive the address of the published metadata queue in the message platform sent by the first data system; Obtain the metadata from the published metadata address according to the address of the published metadata queue.

Citation Information

Patent Citations

  • Apparatus and method for supporting content exchange between different drm domains

    CN101044475A

  • Forest resource heterogeneous data distributed management system

    CN101945126A

  • Metadata Repository and Methods Thereof

    US20090112808A1