A data processing method, apparatus, readable storage medium, and electronic device
By deploying metadata services and metadata replica repositories in cross-datacenter data lakes, the problem of poor performance in cross-datacenter data lakes is solved, enabling rapid response to data query requests and improving data service performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2026-03-10
AI Technical Summary
Data lakes built across multiple data centers tend to have poor performance when providing data services.
Deploy metadata services and build metadata copy repositories in each data center. Through the localized deployment of metadata services and metadata copy repositories, identify target data centers and access the target metadata copy repositories in the same data center to determine metadata information and processing results.
It effectively solves the performance problems caused by cross-data center network latency, can quickly respond to data query requests, and improve the performance of data lake in providing data services.
Smart Images

Figure CN115905256B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a data processing method, apparatus, readable storage medium, and electronic device. Background Technology
[0002] With the explosive growth of data, providing data services for various digital applications through data lakes has become increasingly common. However, when operators' data centers are geographically dispersed, these data lakes are built across multiple data centers, resulting in poor performance when providing data services. Summary of the Invention
[0003] This invention provides a data processing method, apparatus, computer-readable storage medium, and electronic device to solve the technical problem of poor performance of data lakes built across data centers in the prior art when providing data services.
[0004] According to a first aspect of the present invention, a data processing method is provided, applied to a data lake spanning multiple data centers, wherein each data center in the data lake deploys a metadata service, and each data center constructs a metadata replica repository; the method includes:
[0005] Receive data query requests;
[0006] Determine the target data center for executing the data query request;
[0007] The target data center selects the corresponding target metadata service; and based on the data query request, accesses the target metadata copy library belonging to the same data center as the target metadata service to determine the metadata information; and based on the metadata information and the data query request, determines the data processing result.
[0008] Optionally, the cross-datacenter data lake further includes a metadata central repository; after the step of determining the data processing result, the method further includes:
[0009] Based on the data processing results, the target data center accesses the metadata center database and modifies the data in the metadata center database.
[0010] After the data in the metadata central repository is modified, the metadata central repository will synchronize the modification to each metadata replica repository.
[0011] Optionally, the cross-datacenter data lake further includes a metadata service registry; the method further includes:
[0012] The metadata service registration center receives the registration information from each data center when its metadata service is started, and determines the list of metadata services based on the registration information.
[0013] The metadata service registry will distribute the metadata service list to each data center;
[0014] The target data center selects the corresponding target metadata service, including:
[0015] The target data center selects the corresponding target metadata service based on the list of metadata services.
[0016] Optionally, the target data center selects a corresponding target metadata service based on the metadata service list, including:
[0017] The target data center determines the availability status of the metadata services deployed in the target data center based on the metadata service list;
[0018] If the availability status meets the preset conditions, the target data center will select the metadata service deployed in the target data center as the corresponding target metadata service.
[0019] Optionally, the method further includes:
[0020] If the availability status of the target data center does not meet the preset conditions, the target data center determines the parameter information corresponding to the metadata services deployed in other data centers besides the target data center based on the metadata service list.
[0021] The target data center selects the target metadata service with the best parameter information from the metadata services deployed in other data centers.
[0022] The registration information includes data center information; the method further includes:
[0023] If the availability status of the target data center does not meet the preset conditions, the distance between the target data center and other data centers other than the target data center is determined based on the data center information in the metadata service list; the metadata service deployed in the other data center closest to the target data center is determined as the target metadata service.
[0024] According to a second aspect of the present invention, a data processing apparatus is provided, configured in a data lake spanning multiple data centers, wherein each data center in the data lake deploys a metadata service, and each data center has a metadata copy repository; the apparatus includes:
[0025] The request receiving module is used to receive data query requests;
[0026] The center determination module is used to determine the target data center for executing the data query request;
[0027] The data processing module is used to select the corresponding target metadata service in the target data center; and based on the data query request, access the target metadata copy library belonging to the same data center as the target metadata service to determine the metadata information; and based on the metadata information and the data query request, determine the data processing result.
[0028] Optionally, the cross-datacenter data lake further includes a metadata central repository, and the apparatus further includes:
[0029] The change processing module is used by the target data center to access the metadata center database based on the data processing result and perform change processing on the data in the metadata center database.
[0030] The synchronization processing module is used to synchronize the changes made to data in the metadata center database to each metadata copy database after the data in the metadata center database has been changed.
[0031] According to a third aspect of the present invention, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the above-described data processing method.
[0032] According to a fourth aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0033] processor;
[0034] Memory used to store the processor's executable instructions;
[0035] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the data processing method described above.
[0036] Compared with the prior art, the data processing method, apparatus, computer-readable storage medium, and electronic device provided by the present invention have at least the following beneficial effects:
[0037] The technical solution of this invention is applied to a cross-datacenter data lake. It deploys metadata services for each data center within the data lake and builds a data replica library for each data center. Upon receiving a data query request, the data lake determines the target data center to execute the query. This target data center can be any one of at least two data centers included in the data lake. After determining the target data center, the target data center selects the corresponding target metadata service and, based on the data query request, accesses the target metadata replica library belonging to the same data center as the target metadata service to determine the metadata information. The target data center can then determine the data processing result based on the metadata information and the data query request. In the technical solution provided by this invention, the localized deployment of metadata services and metadata replica libraries effectively solves the performance problems caused by cross-datacenter network latency, enabling rapid response to data query requests and improving the performance of the cross-datacenter data lake when providing data services. Attached Figure Description
[0038] To more clearly illustrate the technical solution of this invention, the accompanying drawings used in the description of this invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0039] Figure 1 This is a flowchart illustrating a data processing method provided in an exemplary embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of the system framework of a data processing method provided by an exemplary embodiment of the present invention. Figure 1 ;
[0041] Figure 3 This is a schematic diagram of the system framework of a data processing method provided by an exemplary embodiment of the present invention. Figure 2 ;
[0042] Figure 4 This is a schematic diagram of the structure of a data processing apparatus provided in an exemplary embodiment of the present invention;
[0043] Figure 5 This is a structural diagram of an electronic device provided in an exemplary embodiment of the present invention. Detailed Implementation
[0044] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of these embodiments.
[0045] Exemplary methods
[0046] Figure 1 This is a schematic flowchart of a data processing method provided by an exemplary embodiment of the present invention. The data processing method is applied to a data lake spanning multiple data centers, wherein each data center in the data lake deploys a metadata service and each data center has a metadata copy library.
[0047] The cross-datacenter data lake mentioned in this embodiment refers to a data lake built upon physically distributed data center server rooms. This data lake includes at least two data centers, with the server rooms physically distributed. A metadata service is deployed for each data center, and a metadata replica repository is built for each data center, with the content in each metadata replica repository remaining consistent. This localized deployment of the metadata service and metadata replica repository provides a prerequisite for improving the performance of the data lake. Figure 2 An exemplary data lake framework is shown, which is built from N data centers. Metadata services, namely metadata service A to metadata service N, are deployed in each data center, and a metadata replica library, namely metadata replica library A to metadata replica library N, is built in each data center.
[0048] The method includes at least the following steps:
[0049] Step 10: Receive data query request.
[0050] In one embodiment, the data query request is a structured query statement (SQL statement) used to query data and obtain data processing results.
[0051] Specifically, a data lake includes a unified OLAP (Online Analytical Processing) portal through which data query requests are obtained.
[0052] Step 20: Determine the target data center for executing the data query request.
[0053] The target data center is any one of at least two data centers included in the data lake. In other words, in the technical solution provided in this embodiment, different data centers provide logically unified computing.
[0054] Specifically, an independent computing cluster is deployed in each data center (IDC) room, such as Figure 2The computing clusters A through N shown typically deploy independent computing clusters including Spark, MapReduce, and TEZ engines, and provide logically unified computing capabilities for digital applications through a unified OLAP entry point. Therefore, in this step, determining the target data center for executing the data query request essentially means determining the target computing cluster corresponding to the target data center for executing the data query request.
[0055] In one possible implementation, after receiving a data query request, the unified OLAP entry point randomly selects the target data center from among various data centers. In another possible implementation, data center selection rules are pre-defined, and upon receiving a data query request, the unified OLAP entry point selects the target data center from among various data centers according to these pre-defined rules.
[0056] Step 30: The target data center selects the corresponding target metadata service; and based on the data query request, accesses the target metadata copy library belonging to the same data center as the target metadata service to determine the metadata information; and based on the metadata information and the data query request, determines the data processing result.
[0057] In this step, the target data center selects the corresponding target metadata service. The target metadata service is one of the metadata services deployed in each data center. Since a metadata service is deployed in each data center and a metadata replica library is built, once the target metadata service is determined, it can be determined that the target metadata service belongs to the target metadata replica library of the same data center. Thus, the target data center can directly read the metadata information in the target metadata replica library according to the data query request, and then execute the logic of the data query request based on the metadata information to generate the data processing result.
[0058] In some embodiments, the cross-datacenter data lake further includes a metadata service registry; the method further includes:
[0059] Step 40: The metadata service registration center receives the registration information sent by the metadata service of each data center, and determines the metadata service list based on the registration information.
[0060] In this step, such as Figure 2 As shown, cross-datacenter operations also include a metadata service registry. When a metadata service in each datacenter starts up, it sends registration information to the metadata service registry, i.e., registers with the metadata service registry. The metadata service registry then processes the received registration information to determine the list of metadata services.
[0061] Specifically, a metadata service registry is designed using ZooKeeper (a distributed, open-source distributed application coordination service). Registration information includes, but is not limited to, the metadata service's IP address, port, and the data center information where the metadata service is located. For example,... Figure 2 As shown, metadata service A registers with the metadata service registry, metadata service N registers with the metadata service registry, and the metadata service registry determines that each registered metadata service is included in the metadata service list.
[0062] Step 50: The metadata service registry distributes the metadata service list to each data center.
[0063] In this step, after the metadata service registry determines the list of metadata services, it distributes the list of metadata services to each data center, meaning that each data center will receive the list of metadata services.
[0064] Accordingly, the target data center in step 30 selects the corresponding target metadata service, including:
[0065] Step 301: The target data center selects the corresponding target metadata service based on the metadata service list.
[0066] In this step, the target data center selects the target metadata service based on the received metadata service list, since the metadata service list provides information on all data centers that currently provide metadata services.
[0067] In this embodiment, a metadata service selection mechanism is provided to the target data center through a metadata service registry, thereby realizing distributed metadata services and meeting the performance requirements of large data lakes for metadata services.
[0068] In some embodiments, step 301, where the target data center selects a corresponding target metadata service based on the metadata service list, includes:
[0069] Step 3011: The target data center determines the availability status of the metadata services deployed in the target data center based on the metadata service list.
[0070] In this step, the metadata service list provides the availability status of metadata services deployed in each data center, where availability status indicates whether the metadata service is currently available. The target data center first determines the availability status of the metadata services deployed in the target data center based on the metadata service list. If the target data center corresponds to... Figure 2If computing cluster A is selected, then computing cluster A will prioritize determining whether metadata service A is available.
[0071] Step 3012: If the availability status meets the preset conditions, the target data center selects the metadata service deployed in the target data center as the corresponding target metadata service.
[0072] In this step, pre-set conditions are established. These conditions can be "available" or any parameter indicating that the metadata service deployed in the target data center is available. When the availability status of the target data center meets the pre-set conditions, it indicates that the metadata service deployed in the target data center can provide services normally, and therefore the metadata service deployed in the target data center is selected as the corresponding target metadata service.
[0073] In this embodiment, the target data center first determines the availability status of the metadata service within its own data center. If the metadata service is available, it is identified as the target metadata service. The target data center calls the local metadata service and reads the local metadata copy library, which can effectively reduce the time for cross-data center access, quickly respond to data query requests, and improve the data processing performance of the data lake.
[0074] In some embodiments, the method further includes:
[0075] Step 30113: If the availability status of the target data center does not meet the preset conditions, the target data center determines the parameter information corresponding to the metadata services deployed in other data centers outside the target data center based on the metadata service list.
[0076] Step 30114: The target data center selects the target metadata service with the best parameter information from the metadata services deployed in other data centers.
[0077] In the above steps, considering that the metadata service deployed in the target data center may not be able to provide services normally or its availability does not meet the preset conditions, the target data center determines the parameter information of the metadata services deployed in other data centers based on the metadata service list. This parameter information is used to indicate service performance. Therefore, to ensure the data processing performance of the data lake, the target metadata service with the best parameter information is selected from the metadata services deployed in other data centers. Accessing the target metadata service with the best parameter information by the target data center ensures a faster response time to data query requests. Specifically, a load balancing algorithm can be run to determine the target metadata service when selecting the one with the best parameter information.
[0078] In some embodiments, the registration information includes data center information; the method further includes:
[0079] Step 3015: If the availability status of the target data center does not meet the preset conditions, the target data center determines the distance between the target data center and other data centers other than the target data center based on the data center information in the metadata service list; and determines the metadata service deployed in the other data center closest to the target data center as the target metadata service.
[0080] In this embodiment, if the availability of the metadata service deployed in the target data center does not meet preset conditions, the distance between the target data center and other data centers can be determined based on the data center information in the metadata service list, including data center address information. The metadata service deployed in the other data center closest to the target data center can then be identified as the target metadata service. By reducing the distance between the target data center and the target metadata service, the time for the data lake to respond to data query requests is minimized, thereby improving the data processing performance of the data lake.
[0081] In some embodiments, the cross-datacenter data lake further includes a metadata central repository; after the step of determining the data processing result, the method further includes:
[0082] Step 60: Based on the data processing result, the target data center accesses the metadata center database and performs change processing on the data in the metadata center database.
[0083] In this step, such as Figure 2 As shown, the data lake also includes a metadata central repository. After the target data center generates the data processing results, the target data center accesses the metadata central repository to write the data processing results into the metadata central repository and to perform change processing on the data in the metadata central repository.
[0084] Step 70: After modifying the data in the metadata center database, the metadata center database synchronizes the modification to each metadata copy database.
[0085] In this step, after modifying the data in the metadata central repository, to ensure strong consistency and synchronization between the metadata central repository and the metadata replica repository, the metadata central repository synchronizes the changes to each metadata replica repository. This separation of metadata request read / write is achieved through the metadata central repository and the metadata replica repository, resolving the performance and scalability issues of the metadata database.
[0086] Specifically, such as Figure 2As shown, the metadata service, metadata replica repository, metadata service registry, and metadata center repository in the data lake framework are used to provide unified metadata. Computation clusters A to N in the data lake framework are used to provide unified logical computing. Furthermore, the data lake framework also provides unified logical storage; that is, each data center deploys an independent physical storage cluster, which is implemented through ViewFS to achieve a unified logical cluster. Data access is routed through the logical directory to the physical storage cluster to access the target data file.
[0087] Examples such as Figure 3 As shown, a data lake spanning multiple data centers is... Figure 3 A cross-datacenter data lake is built from N data centers, namely IDC-A, IDC-B, IDC-C, and IDC-N. The application layer sends data query requests to the cross-datacenter data lake. The cross-datacenter data lake receives these requests and selects a target computing cluster from computing clusters A, B, C, to N, such as computing cluster A. Computing cluster A obtains a list of metadata services from the metadata service registry, selecting metadata service A within the same IDC. If metadata service A is unavailable, it selects the nearest metadata service, thus determining the target metadata service. Computing cluster A reads the metadata information of the input tables required by the data query request. The target metadata service accesses the target metadata replica library, which belongs to the same datacenter as the target metadata service, to obtain metadata information. Computing cluster A executes the data query request processing logic and generates data processing results. Computing cluster A accesses the target data table or partition created by the metadata service. The target metadata service accesses the metadata center library to add metadata information to the target data table or partition. The metadata center library performs data change processing and synchronizes the changes to all metadata replica libraries. This enables load balancing and high availability of metadata services through a metadata service registry, allowing the selection of the nearest available metadata service in case of partial service failures. Furthermore, the central metadata database and the metadata replica database maintain strong consistency and synchronization, enabling localized access to the source database, mitigating cross-datacenter latency, improving the performance of cross-datacenter data lakes, and addressing performance scalability issues through diverse data replica databases and service distribution. This results in scalability and high availability for cross-datacenter data lakes.
[0088] The technical solution of this embodiment is applied to a cross-datacenter data lake. It deploys metadata services for each data center in the data lake and builds a data replica library for each data center. Upon receiving a data query request, the data lake determines the target data center to execute the query. This target data center can be any one of at least two data centers included in the data lake. After determining the target data center, the target data center selects the corresponding target metadata service and, based on the data query request, accesses the target metadata replica library belonging to the same data center as the target metadata service to determine the metadata information. Then, the target data center can determine the data processing result based on the metadata information and the data query request. In the technical solution provided by this invention, the localized deployment of metadata services and metadata replica libraries effectively solves the performance problems caused by cross-datacenter network latency, enabling rapid response to data query requests and improving the performance of the cross-datacenter data lake when providing data services.
[0089] Exemplary device
[0090] Based on the same concept as the method embodiments of the present invention, the present invention also provides a data processing apparatus.
[0091] Figure 4 This diagram illustrates the structure of a data processing apparatus provided in an exemplary embodiment of the present invention, configured in a cross-data center data lake, wherein each data center in the data lake deploys a metadata service, and each data center has a metadata replica repository; the apparatus includes:
[0092] Request receiving module 41 is used to receive data query requests;
[0093] Center determination module 42 is used to determine the target data center for executing the data query request;
[0094] The data processing module 43 is used to select the corresponding target metadata service in the target data center; and based on the data query request, access the target metadata copy library that belongs to the same data center as the target metadata service to determine the metadata information; and based on the metadata information and the data query request, determine the data processing result.
[0095] The technical solution of this embodiment is applied to a cross-datacenter data lake. It deploys metadata services for each data center in the data lake and builds a data replica library for each data center. Upon receiving a data query request, the data lake determines the target data center to execute the query. This target data center can be any one of at least two data centers included in the data lake. After determining the target data center, the target data center selects the corresponding target metadata service and, based on the data query request, accesses the target metadata replica library belonging to the same data center as the target metadata service to determine the metadata information. Then, the target data center can determine the data processing result based on the metadata information and the data query request. In the technical solution provided by this invention, the localized deployment of metadata services and metadata replica libraries effectively solves the performance problems caused by cross-datacenter network latency, enabling rapid response to data query requests and improving the performance of the cross-datacenter data lake when providing data services.
[0096] In an exemplary embodiment of the present invention, the cross-datacenter data lake further includes a metadata central repository, and the apparatus further includes:
[0097] The change processing module is used by the target data center to access the metadata center database based on the data processing result and perform change processing on the data in the metadata center database.
[0098] The synchronization processing module is used to synchronize the changes made to data in the metadata center database to each metadata copy database after the data in the metadata center database has been changed.
[0099] In an exemplary embodiment of the present invention, the cross-datacenter data lake further includes a metadata service registry; the apparatus further includes:
[0100] The list determination module is used by the metadata service registry center to receive registration information sent by the metadata service of each data center, and determine the metadata service list based on the registration information; the metadata service registry center then distributes the metadata service list to each data center.
[0101] The data processing module is further configured to select the corresponding target metadata service based on the metadata service list.
[0102] In an exemplary embodiment of the present invention, the data processing module includes:
[0103] A status determination unit is used to determine the availability status of metadata services deployed in the target data center based on the metadata service list.
[0104] The first selection unit is used to select the metadata service deployed in the target data center as the corresponding target metadata service when the availability status meets preset conditions.
[0105] In an exemplary embodiment of the present invention, the apparatus further includes:
[0106] The second selection unit is used to determine the parameter information corresponding to the metadata services deployed in other data centers besides the target data center based on the metadata service list when the available status of the target data center does not meet the preset conditions.
[0107] The target data center selects the target metadata service with the best parameter information from the metadata services deployed in other data centers.
[0108] In an exemplary embodiment of the present invention, the registration information includes data center information; the device further includes:
[0109] The third selection unit is used to determine the distance between the target data center and other data centers outside the target data center based on the data center information in the metadata service list when the availability status of the target data center does not meet the preset conditions; and to determine the metadata service deployed in the other data center closest to the target data center as the target metadata service.
[0110] Exemplary electronic devices
[0111] Figure 5 A block diagram of an electronic device according to an embodiment of the present invention is shown.
[0112] like Figure 5 As shown, the electronic device 50 includes one or more processors 51 and a memory 52.
[0113] The processor 51 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 50 to perform desired functions.
[0114] The memory 52 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 51 may execute the program instructions to implement the data processing methods of the various embodiments of the present invention described above, and / or other desired functions.
[0115] In one example, the electronic device 50 may also include an input device 53 and an output device 54, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0116] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device 50 relevant to the present invention are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 50 may include any other suitable components depending on the specific application.
[0117] Exemplary computer program products and computer-readable storage media
[0118] Sixthly, in addition to the methods and apparatus described above, embodiments of the present invention may also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the data processing methods according to various embodiments of the present invention described in the "Exemplary Methods" section of this specification.
[0119] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0120] Furthermore, embodiments of the present invention may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the data processing methods according to various embodiments of the present invention described in the "Exemplary Methods" section above.
[0121] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0122] The basic principles of the present invention have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in the present invention are merely examples and not limitations, and should not be considered as essential features of each embodiment of the present invention. Furthermore, the specific details of the invention described above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the present invention to the necessity of employing the specific details described above.
[0123] The block diagrams of devices, apparatuses, devices, and systems involved in this invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0124] It should also be noted that in the apparatus, device, and method of the present invention, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of the present invention.
[0125] The above description of aspects of the invention is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features of the invention herein.
[0126] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the invention to the forms described herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A data processing method, characterized by, The application is applied to a cross-data-center data lake, each data center of the data lake is deployed with a metadata service, each data center is built with a metadata copy library, the content in each metadata copy library is consistent, and the cross-data-center data lake further comprises a metadata service registry center; the method comprises: receiving a data query request through a unified online analytical processing portal; the data query request is a structured query statement; determining a target data center for executing the data query request; the target data center selects a corresponding target metadata service from a metadata service list issued by the metadata service registry center; and based on the data query request, accesses a target metadata copy library belonging to the same data center as the target metadata service to determine metadata information; and based on the metadata information and the data query request, determines a data processing result.
2. The method of claim 1, wherein, The cross-data-center data lake further comprises a metadata center library; after the step of determining the data processing result, the method further comprises: the target data center accesses the metadata center library based on the data processing result, and performs change processing on the data in the metadata center library; after the change processing on the data in the metadata center library, the metadata center library synchronizes the change processing to each metadata copy library.
3. The method of claim 1, wherein, The method further comprises: the metadata service registry center receives registration information sent by the metadata service of each data center, and determines a metadata service list based on the registration information; the metadata service registry center issues the metadata service list to each data center; the target data center selecting the corresponding target metadata service comprises: the target data center selects the corresponding target metadata service based on the metadata service list.
4. The method of claim 3, wherein, The target data center selects the corresponding target metadata service based on the metadata service list, comprising: the target data center determines the availability state of the metadata service deployed in the target data center based on the metadata service list; the target data center selects the metadata service deployed in the target data center as the corresponding target metadata service when the availability state meets the preset condition.
5. The method of claim 4, wherein, The method further comprises: when the availability state does not meet the preset condition, the target data center determines the parameter information of the metadata service deployed in other data centers except the target data center based on the metadata service list; the target data center selects the target metadata service with the best parameter information from the metadata services deployed in other data centers.
6. The method of claim 4, wherein, The registration information comprises machine room information; the method further comprises: when the availability state does not meet the preset condition, the target data center determines the distance between the target data center and other data centers except the target data center based on the machine room information in the metadata service list; and determines the metadata service deployed in the other data center closest to the target data center as the target metadata service.
7. A data processing apparatus, characterized by, A data lake across data centers is configured, a metadata service is deployed in each data center, a metadata copy library is built in each data center, the content in each metadata copy library is consistent, and a metadata service registration center is further included in the data lake across data centers; the device comprises: A request receiving module is configured to receive a data query request through a unified online analytical processing portal; the data query request is a structured query statement; A center determining module is configured to determine a target data center for executing the data query request; A data processing module is configured to select a corresponding target metadata service from a metadata service list issued by the metadata service registration center in the target data center; access a target metadata copy library belonging to a same data center as the target metadata service based on the data query request, and determine metadata information; and determine a data processing result based on the metadata information and the data query request.
8. The apparatus of claim 7, wherein, The data lake across data centers further comprises a metadata center library, and the device further comprises: A change processing module is configured to access the metadata center library based on the data processing result in the target data center, and perform change processing on data in the metadata center library; A synchronization processing module is configured to synchronize change processing to each metadata copy library after the data in the metadata center library is processed. 9.A computer readable storage medium, the storage medium storing a computer program, the computer program being used to execute the data processing method in any one of claims 1-6. 10.An electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the data processing method in any one of claims 1-6.
Citation Information
Patent Citations
Multi- data-centre hadoop distributed file system (HDFS) data read-write system and method
CN104113597A
Cross-regional data scheduling method and device, equipment and storage medium
CN115292280A