A data synchronization method, apparatus, system, device, and medium for a data lake
By pausing the incremental data synchronization of the target data lake table in the data lake message queue, the stranded data is in the message queue, and the stranded synchronization is performed after the existing synchronization is completed, the problem of low data synchronization efficiency in the existing technology is solved, and efficient data synchronization processing is achieved.
Patent Information
- Application Number
- CN202310817388.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-07-05
AI Technical Summary
When synchronizing the stock data of the target data lake table, the data consumption of the entire data lake message queue is usually necessary, resulting in the data synchronization of other data lake tables being affected, reducing the overall data synchronization efficiency.
By pausing the incremental data synchronization of only the target data lake table in the data lake message queue, the stranded data is in the message queue, and after the inventory synchronization is completed, the stranded data is synchronized to the target data lake table and the warehouse through the stranded synchronization service, and the incremental synchronization service is restarted.
When synchronizing the stock data of the target data lake table, it does not affect the data synchronization of other data lake tables, reducing the complexity of data management and data processing difficulty, and improving data synchronization efficiency.
Smart Images

Figure CN116932645B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data synchronization processing, and particularly to a data synchronization method, device, system, equipment and medium for a data lake. Background Art
[0002] In the traditional data middle platform architecture, the data lake, as a buffer zone between the original business data source and the data warehouse, provides the ability to collect and converge different original business data, and at the same time provides a basic data source for the upper-layer data warehouse modeling. With the development of data synchronization technology, especially the change data capture technology based on database logs, real-time data collection and synchronization solutions have gradually become the mainstream in the industry.
[0003] Based on this, in the construction of the middle platform, different original business data are first collected in real time into the data lake, and the upper-layer data warehouse selects the basic data tables that need to be synchronized from the data lake according to the requirements of the data warehouse model for real-time data synchronization processing. To ensure data integrity, the synchronization processing of the data lake data by the upper-layer data warehouse at least includes: 1. Synchronizing the stock data already in the target data lake table; 2. Synchronizing the incremental data corresponding to the target data lake table in the data lake message queue collected in real time.
[0004] In the existing technical solutions, when the target data warehouse newly adds a data synchronization requirement for the data lake table, in order to synchronize the stock data in the target data lake table, a common approach is to suspend the data consumption of the entire data lake message queue to achieve the purpose of suspending the incremental data corresponding to the target data lake table from entering the lake. However, since the original business data collected in real time in the data lake message queue is not only the incremental data corresponding to the target data lake table, but also involves the incremental data corresponding to other data lake tables, directly suspending the data consumption of the entire data lake message queue will affect the data synchronization of other tables in the data lake, thus reducing the overall data synchronization efficiency; another common approach is to temporarily store the consumed incremental data corresponding to the target data lake table in a temporary storage medium, such as Hive (a data warehouse tool based on the distributed system infrastructure), and wait until the stock data in the target data lake table is synchronized, and then perform lake entry and synchronization processing on the incremental data in the temporary storage medium. Although this solution only suspends the lake entry of the incremental data corresponding to the target data lake table in the data lake message queue when synchronizing the stock data in the target data lake table, since the incremental data corresponding to the target data lake table needs to be temporarily stored in a temporary storage medium, it increases the complexity of temporary data management and the difficulty of connecting the subsequent temporary data processing and real-time incremental data processing, and is prone to problems of unsmooth connection, thereby reducing the data processing volume of the data lake message queue per unit time, and thus reducing the overall data synchronization efficiency. Summary of the Invention
[0005] This application provides a data synchronization method, apparatus, system, device and medium for a data lake to overcome the defects of the above-mentioned existing technologies, which can improve the data synchronization efficiency on the basis of reducing the complexity of data management and the difficulty of data processing connection.
[0006] To solve the above technical problems, this application provides the following technical solutions:
[0007] According to the first aspect of the embodiments of this application, a data synchronization method for a data lake is provided, including:
[0008] An external data warehouse sends a first data synchronization request to the data lake message queue; the first data synchronization request carries first data synchronization relationship information;
[0009] The incremental synchronization service obtains the first data synchronization request from the data lake message queue;
[0010] Based on the first data synchronization request, the incremental synchronization service sends the first data synchronization relationship information to the global synchronization status cache to update the synchronization relationship in the global synchronization status cache, obtaining a first target synchronization relationship; and based on the first data synchronization request, the incremental synchronization service sends a full synchronization request to the full synchronization service, so that the full synchronization service synchronizes the full data in the first target data lake table to the first target data warehouse corresponding to the first target data lake table; the full synchronization request carries full synchronization relationship information, and the full synchronization relationship information is determined based on the first data synchronization relationship information; and based on the full synchronization request, the incremental synchronization service suspends the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue, so that the incremental data corresponding to the first target data lake table remains in the data lake message queue;
[0011] When the incremental synchronization service queries that the synchronization progress information in the global synchronization status cache indicates that the full synchronization task corresponding to the first target data lake table has been completed, it sends a retention synchronization request to the retention synchronization service, so that the retention synchronization service sends the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the retention synchronization request carries retention synchronization relationship information, and the retention synchronization relationship information is determined based on the first data synchronization relationship information; and the incremental synchronization service is suspended;
[0012] When the synchronization progress information indicates that all the stranded synchronization tasks corresponding to the first target data lake table have been completed, the incremental synchronization service is restarted, and the incremental synchronization service sends the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse.
[0013] Further, the method includes two data synchronization scenarios, specifically including the incremental synchronization scenario in the steady state and the initialization synchronization scenario in the non-steady state. The step from the external data warehouse sending the first data synchronization request to the data lake message queue to the incremental synchronization service sending the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizing it to the first target data warehouse is determined based on the initialization synchronization scenario in the non-steady state. The data synchronization steps in the incremental synchronization scenario in the steady state include:
[0014] The external data warehouse sends a second data synchronization request to the data lake message queue; the second data synchronization request carries second data synchronization relationship information;
[0015] The incremental synchronization service obtains the second data synchronization request from the data lake message queue;
[0016] The incremental synchronization service sends the second data synchronization relationship information to the global synchronization status cache based on the second data synchronization request to update the synchronization relationship in the global synchronization status cache and obtain a second target synchronization relationship;
[0017] The incremental synchronization service sends the incremental data corresponding to the second target data lake table newly added in the data lake message queue to the second target data lake table based on the second target synchronization relationship, or further synchronizes it to the second target data warehouse if there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse.
[0018] Further, the method also includes the step of determining the incremental data corresponding to the stranded first target data lake table. Determining the incremental data corresponding to the stranded first target data lake table includes:
[0019] When the number of first stock synchronization tasks corresponding to the first target data lake table in the synchronization progress information queried by the incremental synchronization service is zero, the incremental synchronization service obtains the first offset position of the incremental data corresponding to the first target data lake table in the data lake message queue, and records the first offset position in the global synchronization status cache; the number of first stock synchronization tasks is the number of stock synchronization tasks corresponding to the first target data lake table recorded in the global synchronization status cache before the incremental synchronization service obtains the first data synchronization request;
[0020] When the number of second stock synchronization tasks corresponding to the first target data lake table in the synchronization progress information queried by the incremental synchronization service is zero, the incremental synchronization service obtains the second offset position of the incremental data corresponding to the first target data lake table in the data lake message queue, and records the second offset position in the global synchronization status cache; the number of second stock synchronization tasks is the number of stock synchronization tasks corresponding to the first target data lake table recorded in the global synchronization status cache after the incremental synchronization service obtains the first data synchronization request;
[0021] The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache.
[0022] Further, the stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache, including:
[0023] The incremental synchronization service obtains the first offset position and the second offset position from the global synchronization status cache;
[0024] The incremental synchronization service sends the stranded synchronization request to the stranded synchronization service; the stranded synchronization request also carries the first offset position and the second offset position;
[0025] The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the stranded synchronization request.
[0026] Further, the method further includes:
[0027] When the number of first stock synchronization tasks is not zero, the incremental synchronization service does not obtain the first offset position.
[0028] Further, the method further includes:
[0029] When the number of the second inventory synchronization tasks is not zero, the incremental synchronization service does not obtain the position of the second offset.
[0030] Further, the method further includes a step of determining that all the inventory synchronization tasks corresponding to the first target data lake table have been completed. Determining that all the inventory synchronization tasks corresponding to the first target data lake table have been completed includes:
[0031] After completing one of the inventory synchronization tasks, the inventory synchronization service sends a first indication message to the data lake message queue;
[0032] The data lake message queue receives the first indication message;
[0033] The incremental synchronization service obtains the first indication message from the data lake message queue;
[0034] The incremental synchronization service sends the first indication message to the global synchronization status cache;
[0035] The global synchronization status cache updates the synchronization progress information based on the first indication message;
[0036] Repeat the steps from completing one of the inventory synchronization tasks to the global synchronization status cache updating the synchronization progress information based on the first indication message until the synchronization progress information indicates that the number of the inventory synchronization tasks corresponding to the first target data lake table is zero, then it is determined that all the inventory synchronization tasks corresponding to the first target data lake table have been completed.
[0037] According to a second aspect of the embodiments of the present application, there is provided a data synchronization device for a data lake. The device includes:
[0038] A synchronization request sending module, configured to send a first data synchronization request to the data lake message queue by an external data warehouse; the first data synchronization request carries first data synchronization relationship information;
[0039] A synchronization request obtaining module, configured to obtain the first data synchronization request from the data lake message queue by the incremental synchronization service;
[0040] The stock synchronization module is used for the incremental synchronization service to send the first data synchronization relationship information to the global synchronization status cache based on the first data synchronization request, so as to update the synchronization relationship in the global synchronization status cache and obtain the first target synchronization relationship; and for the incremental synchronization service to send a stock synchronization request to the stock synchronization service based on the first data synchronization request, so that the stock synchronization service synchronizes the stock data in the first target data lake table to the first target data warehouse corresponding to the first target data lake table; the stock synchronization request carries stock synchronization relationship information, and the stock synchronization relationship information is determined based on the first data synchronization relationship information; and for the incremental synchronization service to pause the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue, so that the incremental data corresponding to the first target data lake table stays in the data lake message queue;
[0041] The retention synchronization module is used for the incremental synchronization service to send a retention synchronization request to the retention synchronization service when the synchronization progress information in the global synchronization status cache indicates that the stock synchronization task corresponding to the first target data lake table has been completely completed, so that the retention synchronization service sends the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the retention synchronization request carries retention synchronization relationship information, and the retention synchronization relationship information is determined based on the first data synchronization relationship information; and for pausing the incremental synchronization service;
[0042] The incremental synchronization module is used for the incremental synchronization service to be restarted when the synchronization progress information indicates that the retention synchronization task corresponding to the first target data lake table has been completely completed, and the incremental synchronization service sends the newly added incremental data corresponding to the first target data lake table in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse.
[0043] According to the third aspect of the embodiments of the present application, an electronic device is provided, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the data synchronization method of the data lake described in any one of the above.
[0044] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one instruction or at least one program segment is stored in the storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the data synchronization method of the data lake described in any one of the above.
[0045] With the above technical solution, the present application has the following beneficial effects:
[0046] A data synchronization method, device, system, equipment and medium for a data lake provided by the present application. In this technical solution, when the incremental synchronization service receives a first data synchronization request, the stock synchronization service is started to synchronize the stock data in the first target data lake table to the first target data warehouse; and the incremental data corresponding to the first target data lake table is retained in the data lake message queue; when all the stock synchronization tasks are completed, the retention synchronization service is started to send the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronize it to the first target data warehouse; and the incremental synchronization service is paused; when all the retention synchronization tasks are completed, the incremental synchronization service is restarted, and the incremental synchronization service sends the incremental data corresponding to the newly added first target data lake table in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse. Through this technical solution, during the data synchronization process, only the inflow of the incremental data corresponding to the first target data lake table can be paused, and the inflow of data in other data lake tables is not affected, and the smooth connection between the retention data synchronization and the incremental data synchronization can be ensured, that is, the data synchronization efficiency can be improved on the basis of reducing the complexity of data management and the difficulty of data processing connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is a schematic flowchart of a data synchronization method for a data lake provided by an embodiment of the present application;
[0049] Figure 2 It is a structural block diagram of a data synchronization device for a data lake provided by an embodiment of the present application;
[0050] Figure 3 It is a schematic flowchart of a data synchronization system for a data lake provided by an embodiment of the present application;
[0051] Figure 4 It is a schematic flowchart of the incremental synchronization of a data synchronization system for a data lake provided by an embodiment of the present application;
[0052] Figure 5 It is a schematic flowchart of the stock synchronization of a data synchronization system for a data lake provided by an embodiment of the present application;
[0053] Figure 6 This is a schematic flowchart of the retention synchronization of a data synchronization system for a data lake provided by an embodiment of the present application;
[0054] Figure 7 This is a hardware structure block diagram of an electronic device for running a data synchronization method for a data lake provided by an embodiment of the present application. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0056] As used herein, the term "one embodiment" or "embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present application. In the description of the embodiments of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "top", "bottom", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Moreover, the terms "first", "second", etc. are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0057] Please refer to Figure 1, which shows a schematic flowchart of a data synchronization method for a data lake provided by an embodiment of the present application. This specification provides method operation steps as described in the embodiments or flowcharts. However, based on routine or non-creative labor, more or fewer operation steps may be included. The step sequences listed in the embodiments are only one of the ways of the execution sequences of numerous steps and do not represent the only execution sequence. When the actual system or server product executes, it can be executed in the method sequence shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). The method may include:
[0058] Step S101: The external data warehouse sends a first data synchronization request to the data lake message queue; the first data synchronization request carries first data synchronization relationship information;
[0059] Step S102: The incremental synchronization service obtains the first data synchronization request from the data lake message queue;
[0060] Step S103: Based on the first data synchronization request, the incremental synchronization service sends the first data synchronization relationship information to the global synchronization status cache to update the synchronization relationship in the global synchronization status cache, obtaining a first target synchronization relationship; and based on the first data synchronization request, the incremental synchronization service sends a stock synchronization request to the stock synchronization service, so that the stock synchronization service synchronizes the stock data in the first target data lake table to the first target data warehouse corresponding to the first target data lake table; the stock synchronization request carries stock synchronization relationship information, and the stock synchronization relationship information is determined based on the first data synchronization relationship information; and based on the stock synchronization request, the incremental synchronization service suspends the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue, so that the incremental data corresponding to the first target data lake table stays in the data lake message queue;
[0061] Step S104: When the incremental synchronization service queries that the synchronization progress information in the global synchronization status cache indicates that the stock synchronization task corresponding to the first target data lake table has been completely completed, it sends a retention synchronization request to the retention synchronization service, so that the retention synchronization service sends the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the retention synchronization request carries retention synchronization relationship information, and the retention synchronization relationship information is determined based on the first data synchronization relationship information; and the incremental synchronization service is suspended;
[0062] Step S105: When the synchronization progress information indicates that all the remaining synchronization tasks corresponding to the first target data lake table have been completed, restart the incremental synchronization service, and the incremental synchronization service sends the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse;
[0063] Among them, the synchronization progress information is determined based on the stock synchronization service, the remaining synchronization service, and the incremental synchronization service.
[0064] In a specific embodiment, through step S101, the external data warehouse sends a first data synchronization request to the data lake message queue; where the external data warehouse is an external data warehouse system; through step S102, the incremental synchronization service obtains the first data synchronization request from the data lake message queue; through step S103, update the synchronization relationship in the global synchronization status cache, and the incremental synchronization service sends a stock synchronization request to the stock synchronization service through the stock synchronization message queue, so that the stock synchronization service synchronizes the stock data in the first target data lake table to the data warehouse message queue based on the stock synchronization request, and then the data warehouse message queue transmits it to the first target data warehouse to complete the stock data synchronization. During the stock data synchronization process, the incremental data corresponding to the first target data lake table is paused from entering the lake and remains in the data lake message queue; through step S104, when the stock data synchronization is all completed, start the remaining synchronization service to synchronize the data remaining in the data lake message queue. At this time, the incremental synchronization service sends a remaining synchronization request to the remaining synchronization service through the remaining task message queue, so that the remaining synchronization service consumes the remaining data corresponding to the first target data lake table based on the remaining synchronization request, and after consumption, it enters the first target data lake table, and further synchronizes the remaining data from the first target data lake table to the data warehouse message queue, and then the data warehouse message queue transmits it to the first target data warehouse, and the incremental synchronization service is paused during the remaining data synchronization process; through step S105, when the remaining data synchronization is all completed, restart the incremental synchronization service, and the incremental synchronization service consumes the newly added incremental data corresponding to the first target data lake table, and after consumption, it enters the first target data lake table, and further synchronizes the newly added incremental data from the first target data lake table to the data warehouse message queue, and then the data warehouse message queue transmits it to the first target data warehouse.
[0065] The data synchronization method for a data lake described in this embodiment can, during the data synchronization process, only pause the incremental data corresponding to the first target data lake table from entering the lake, and the data entry of other data lake tables is not affected, and it can ensure the smooth connection between the remaining data synchronization and the incremental data synchronization, that is, it can improve the data synchronization efficiency on the basis of reducing the data management complexity and the data processing connection difficulty.
[0066] In an optional implementation, the method includes two data synchronization scenarios, specifically including an incremental synchronization scenario in a steady state and an initialization synchronization scenario in an unsteady state. The step in which the external data warehouse sends a first data synchronization request to the data lake message queue and the incremental synchronization service sends the incremental data corresponding to the newly added first target data lake table in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse is determined based on the initialization synchronization scenario in the unsteady state. The data synchronization steps of the incremental synchronization scenario in the steady state include:
[0067] The external data warehouse sends a second data synchronization request to the data lake message queue; the second data synchronization request carries second data synchronization relationship information;
[0068] The incremental synchronization service obtains the second data synchronization request from the data lake message queue;
[0069] The incremental synchronization service sends the second data synchronization relationship information to the global synchronization status cache based on the second data synchronization request to update the synchronization relationship in the global synchronization status cache and obtain a second target synchronization relationship;
[0070] The incremental synchronization service sends the incremental data corresponding to the newly added second target data lake table in the data lake message queue to the second target data lake table based on the second target synchronization relationship, or further synchronizes it to the second target data warehouse if there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse.
[0071] Specifically, when the data synchronization requirement of the external data warehouse system for the data lake data does not change and the status of the incremental synchronization task in the global synchronization status cache is the "running" state, that is, when the external data warehouse sends a second data synchronization request, the data synchronization scenario is the incremental synchronization scenario in the steady state; when the external data warehouse newly adds a data synchronization requirement for the data lake table, that is, when the external data warehouse sends a first data synchronization request, the data synchronization scenario is the initialization synchronization scenario in the unsteady state;
[0072] In the incremental synchronization scenario of the steady state, when there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse, the incremental data corresponding to the newly added second target data lake table is consumed through the incremental synchronization service. After consumption, it is ingested into the second target data lake table, and further, the newly added incremental data is synchronized from the second target data lake table to the data warehouse message queue, and then transmitted from the data warehouse message queue to the second target data warehouse; when there is no corresponding synchronization relationship between the second target data lake table and the second target data warehouse, the incremental data corresponding to the newly added second target data lake table is consumed through the incremental synchronization service. After consumption, it is ingested into the second target data lake table, and the newly added incremental data is not synchronized from the second target data lake table to the data warehouse message queue.
[0073] In an optional implementation manner, the method further includes the step of determining the incremental data corresponding to the stranded first target data lake table. Determining the incremental data corresponding to the stranded first target data lake table includes:
[0074] When the first inventory synchronization task number corresponding to the first target data lake table in the synchronization progress information queried by the incremental synchronization service is zero, the incremental synchronization service obtains the first offset position of the incremental data corresponding to the first target data lake table in the data lake message queue, and records the first offset position in the global synchronization status cache; the first inventory synchronization task number is the inventory synchronization task number corresponding to the first target data lake table recorded in the global synchronization status cache before the incremental synchronization service obtains the first data synchronization request.
[0075] When the second inventory synchronization task number corresponding to the first target data lake table in the synchronization progress information queried by the incremental synchronization service is zero, the incremental synchronization service obtains the second offset position of the incremental data corresponding to the first target data lake table in the data lake message queue, and records the second offset position in the global synchronization status cache; the second inventory synchronization task number is the inventory synchronization task number corresponding to the first target data lake table recorded in the global synchronization status cache after the incremental synchronization service obtains the first data synchronization request.
[0076] The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache.
[0077] Specifically, after the incremental synchronization service obtains the first data synchronization request, it updates the number of stock synchronization tasks recorded in the global synchronization status cache based on the first data synchronization request. The number of stock synchronization tasks before the update is the first stock synchronization task number, and the number of stock synchronization tasks after the update is the second stock synchronization task number.
[0078] In an optional implementation, the stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache, including:
[0079] The incremental synchronization service obtains the first offset position and the second offset position from the global synchronization status cache;
[0080] The incremental synchronization service sends the stranded synchronization request to the stranded synchronization service; the first offset position and the second offset position are also carried in the stranded synchronization request;
[0081] The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the stranded synchronization request.
[0082] Specifically, the first offset position is begin_offset, and the second offset position is end_offset.
[0083] In an optional implementation, the method further includes:
[0084] When the first stock synchronization task number is not zero, the incremental synchronization service does not obtain the first offset position.
[0085] In an optional implementation, the method further includes:
[0086] When the second stock synchronization task number is not zero, the incremental synchronization service does not obtain the second offset position.
[0087] In an optional implementation, the method further includes the step of determining that all the stock synchronization tasks corresponding to the first target data lake table have been completed. Determining that all the stock synchronization tasks corresponding to the first target data lake table have been completed includes:
[0088] After completing one stock synchronization task, the stock synchronization service sends a first indication message to the data lake message queue;
[0089] The data lake message queue receives the first indication message;
[0090] The incremental synchronization service obtains the first indication information from the data lake message queue;
[0091] The incremental synchronization service sends the first indication information to the global synchronization status cache;
[0092] The global synchronization status cache updates the synchronization progress information based on the first indication information;
[0093] Repeat the step of completing one stock synchronization task until the global synchronization status cache updates the synchronization progress information based on the first indication information, until the synchronization progress information indicates that the number of stock synchronization tasks corresponding to the first target data lake table is zero, then it is determined that all the stock synchronization tasks corresponding to the first target data lake table have been completed.
[0094] Specifically, for each completed stock synchronization task, the number of stock synchronization tasks recorded in the global synchronization status cache decreases by 1.
[0095] In an optional embodiment, the method further includes the step of restarting the incremental synchronization service. Restarting the incremental synchronization service includes:
[0096] After completing the stranded synchronization task, the stranded synchronization service sends second indication information to the global synchronization status cache;
[0097] The global synchronization status cache receives the second indication information;
[0098] Based on the second indication information, the global synchronization status cache changes the status of the incremental synchronization service from paused to running, so that the incremental synchronization service is restarted.
[0099] In an optional embodiment, the data in the data lake message queue is collected based on the original business data.
[0100] Specifically, the real-time incremental collection of the original business data is realized through the Change Data Capture (CDC) technology.
[0101] As can be seen from the above technical solutions of the embodiments of the present application, in the embodiments of the present application, when the incremental synchronization service obtains the first data synchronization request, the stock synchronization service is started to synchronize the stock data in the first target data lake table to the first target data warehouse; and the incremental data corresponding to the first target data lake table is retained in the data lake message queue; when all the stock synchronization tasks are completed, the retention synchronization service is started to send the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronize it to the first target data warehouse; and the incremental synchronization service is paused; when all the retention synchronization tasks are completed, the incremental synchronization service is restarted, and the incremental synchronization service sends the incremental data corresponding to the newly added first target data lake table in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse. Through this technical solution, during the data synchronization process, only the incremental data entry of the first target data lake table can be paused, and the data entry of other data lake tables is not affected, and the smooth connection between the retention data synchronization and the incremental data synchronization can be ensured, that is, the data synchronization efficiency can be improved on the basis of reducing the data management complexity and the data processing connection difficulty.
[0102] Corresponding to the data synchronization method of the data lake provided in the above embodiment, the embodiment of the present application also provides a data synchronization device for the data lake. Since the data synchronization device for the data lake provided in the embodiment of the present application corresponds to the data synchronization method of the data lake provided in the above embodiment, the implementation manners of the foregoing data synchronization method of the data lake are also applicable to the data synchronization device for the data lake provided in this embodiment, and will not be described in detail in this embodiment.
[0103] Please refer to Figure 2 , which shows a structural block diagram of a data synchronization device for a data lake provided by an embodiment of the present application. The device includes:
[0104] 001: Synchronization request sending module, configured to send a first data synchronization request from an external data warehouse to the data lake message queue; the first data synchronization request carries first data synchronization relationship information;
[0105] 002: Synchronization request acquisition module, configured to acquire the first data synchronization request from the data lake message queue by the incremental synchronization service;
[0106] 003: Stock synchronization module, which is used for the incremental synchronization service to send the first data synchronization relationship information to the global synchronization status cache based on the first data synchronization request, so as to update the synchronization relationship in the global synchronization status cache and obtain the first target synchronization relationship; and is used for the incremental synchronization service to send a stock synchronization request to the stock synchronization service based on the first data synchronization request, so that the stock synchronization service synchronizes the stock data in the first target data lake table to the first target data warehouse corresponding to the first target data lake table; the stock synchronization request carries stock synchronization relationship information, and the stock synchronization relationship information is determined based on the first data synchronization relationship information; and is used for the incremental synchronization service to pause the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue, so that the incremental data corresponding to the first target data lake table stays in the data lake message queue;
[0107] 004: Stagnant synchronization module, which is used for the incremental synchronization service to send a stagnant synchronization request to the stagnant synchronization service when the synchronization progress information in the global synchronization status cache indicates that the stock synchronization task corresponding to the first target data lake table has been completed, so that the stagnant synchronization service sends the incremental data corresponding to the first target data lake table stagnated in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the stagnant synchronization request carries stagnant synchronization relationship information, and the stagnant synchronization relationship information is determined based on the first data synchronization relationship information; and is used for pausing the incremental synchronization service;
[0108] 005: Incremental synchronization module, which is used for the incremental synchronization service to restart when the synchronization progress information indicates that the stagnant synchronization task corresponding to the first target data lake table has been completed, and the incremental synchronization service sends the newly added incremental data corresponding to the first target data lake table in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse.
[0109] In a specific embodiment, through the synchronous request sending module, the external data warehouse sends a first data synchronization request to the data lake message queue; wherein, the external data warehouse is an external data warehouse system; through the synchronous request acquisition module, the incremental synchronization service acquires the first data synchronization request from the data lake message queue; through the stock synchronization module, the synchronization relationship in the global synchronization status cache is updated, and the incremental synchronization service sends a stock synchronization request to the stock synchronization service through the stock synchronization message queue, so that the stock synchronization service synchronizes the stock data in the first target data lake table to the data warehouse message queue based on the stock synchronization request, and then the data warehouse message queue transmits it to the first target data warehouse to complete the stock data synchronization, and during the stock data synchronization process, the inflow of incremental data corresponding to the first target data lake table is paused, making it stay in the data lake message queue; through the retention synchronization module, when the stock data synchronization is all completed, the retention synchronization service is started to synchronize the data staying in the data lake message queue. At this time, the incremental synchronization service sends a retention synchronization request to the retention synchronization service through the retention task message queue, so that the retention synchronization service consumes the retention data corresponding to the first target data lake table based on the retention synchronization request, and after consumption, it is ingested into the first target data lake table, and further synchronizes the retention data from the first target data lake table to the data warehouse message queue, and then the data warehouse message queue transmits it to the first target data warehouse, and the incremental synchronization service is paused during the retention data synchronization process; through the incremental synchronization module, when the retention data synchronization is all completed, the incremental synchronization service is restarted, and the incremental data corresponding to the newly added first target data lake table is consumed through the incremental synchronization service, and after consumption, it is ingested into the first target data lake table, and further synchronizes the newly added incremental data from the first target data lake table to the data warehouse message queue, and then the data warehouse message queue transmits it to the first target data warehouse.
[0110] In an alternative embodiment, the apparatus further includes:
[0111] An incremental synchronization request sending module, configured to enable the external data warehouse to send a second data synchronization request to the data lake message queue; the second data synchronization request carries second data synchronization relationship information;
[0112] An incremental synchronization request acquisition module, configured to enable the incremental synchronization service to acquire the second data synchronization request from the data lake message queue;
[0113] An incremental synchronization relationship update module, configured to enable the incremental synchronization service to send the second data synchronization relationship information to the global synchronization status cache based on the second data synchronization request to update the synchronization relationship in the global synchronization status cache to obtain a second target synchronization relationship;
[0114] An incremental lake input and synchronization module is used for the incremental synchronization service to send the incremental data corresponding to the second target data lake table newly added in the data lake message queue to the second target data lake table based on the second target synchronization relationship, or further synchronize it to the second target data warehouse in the case where there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse.
[0115] Specifically, the device can implement two data synchronization scenarios, specifically including a stable-state incremental synchronization scenario and an unstable-state initialization synchronization scenario. Modules 001 to 005 are modules for implementing the unstable-state initialization synchronization scenario, and the incremental synchronization request sending module, the incremental synchronization request obtaining module, the incremental synchronization relationship updating module, and the incremental lake input and synchronization module are modules for implementing the stable-state incremental synchronization scenario; when the data synchronization requirement of the external data warehouse system for the data lake data does not change, and in the global synchronization status cache, the incremental synchronization task status is in the "running" state, that is, in the case where the external data warehouse sends a second data synchronization request, the data synchronization scenario is a stable-state incremental synchronization scenario; when the external data warehouse newly adds a data synchronization requirement for a data lake table, that is, in the case where the external data warehouse sends a first data synchronization request, the data synchronization scenario is an unstable-state initialization synchronization scenario;
[0116] In the stable-state incremental synchronization scenario, in the case where there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse, the incremental data corresponding to the newly added second target data lake table is consumed through the incremental synchronization service, and after consumption, it is input into the second target data lake table, and further, the newly added incremental data is synchronized from the second target data lake table to the data warehouse message queue, and then transmitted from the data warehouse message queue to the second target data warehouse; in the case where there is no corresponding synchronization relationship between the second target data lake table and the second target data warehouse, the incremental data corresponding to the newly added second target data lake table is consumed through the incremental synchronization service, and after consumption, it is input into the second target data lake table, and the newly added incremental data is not synchronized from the second target data lake table to the data warehouse message queue.
[0117] In an optional implementation manner, the device further includes:
[0118] A first judgment module, which is used to judge whether the first stock synchronization task number is zero;
[0119] The first recording module is configured to, when the incremental synchronization service queries that the number of first stock synchronization tasks corresponding to the first target data lake table in the synchronization progress information is zero, the incremental synchronization service obtains the first offset position of the incremental data corresponding to the first target data lake table in the data lake message queue and records the first offset position in the global synchronization status cache; the number of first stock synchronization tasks is the number of stock synchronization tasks corresponding to the first target data lake table recorded in the global synchronization status cache before the incremental synchronization service obtains the first data synchronization request; and is configured to, when the number of first stock synchronization tasks is not zero, the incremental synchronization service does not obtain the first offset position;
[0120] The second judgment module is configured to judge whether the number of second stock synchronization tasks is zero;
[0121] The second recording module is configured to, when the incremental synchronization service queries that the number of second stock synchronization tasks corresponding to the first target data lake table in the synchronization progress information is zero, the incremental synchronization service obtains the second offset position of the incremental data corresponding to the first target data lake table in the data lake message queue and records the second offset position in the global synchronization status cache; the number of second stock synchronization tasks is the number of stock synchronization tasks corresponding to the first target data lake table recorded in the global synchronization status cache after the incremental synchronization service obtains the first data synchronization request; and is configured to, when the number of second stock synchronization tasks is not zero, the incremental synchronization service does not obtain the second offset position;
[0122] The stranded data determination module is configured to, based on the first offset position and the second offset position in the global synchronization status cache, the stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table.
[0123] Specifically, after the incremental synchronization service obtains the first data synchronization request, it updates the number of stock synchronization tasks recorded in the global synchronization status cache based on the first data synchronization request. The number of stock synchronization tasks before the update is the number of first stock synchronization tasks, and the number of stock synchronization tasks after the update is the number of second stock synchronization tasks.
[0124] In an alternative embodiment, the above-mentioned stranded data determination module may include:
[0125] The position acquisition module is configured to, for the incremental synchronization service, obtain the first offset position and the second offset position from the global synchronization status cache;
[0126] An information sending module, configured to send the retention synchronization request to the retention synchronization service by the incremental synchronization service; the first offset position and the second offset position are also carried in the retention synchronization request;
[0127] A data determination module, configured to determine, by the retention synchronization service based on the retention synchronization request, the incremental data corresponding to the first target data lake table of the retention.
[0128] Specifically, the first offset position is begin_offset, and the second offset position is end_offset.
[0129] In an optional implementation manner, the device further includes:
[0130] A first indication sending module, configured to, when completing a stock synchronization task, send a first indication message by the stock synchronization service to the data lake message queue;
[0131] A first indication receiving module, configured to receive the first indication message by the data lake message queue;
[0132] A first indication obtaining module, configured to obtain the first indication message from the data lake message queue by the incremental synchronization service;
[0133] A first indication resending module, configured to send the first indication message to the global synchronization status cache by the incremental synchronization service;
[0134] An update module, configured to update the synchronization progress information by the global synchronization status cache based on the first indication message;
[0135] A repetition module, configured to repeatedly execute the first indication sending module, the first indication receiving module, the first indication obtaining module, the first indication resending module, and the update module until the synchronization progress information indicates that the number of stock synchronization tasks corresponding to the first target data lake table is zero, then determine that the stock synchronization tasks corresponding to the first target data lake table have all been completed, and end the repeated execution.
[0136] Specifically, for each completed stock synchronization task, the number of stock synchronization tasks recorded in the global synchronization status cache is decremented by 1.
[0137] In an optional implementation manner, the device further includes:
[0138] A second indication sending module, configured to send a second indication message to the global synchronization status cache by the retention synchronization service when completing the retention synchronization task;
[0139] A second indication receiving module, configured to receive the second indication information by the global synchronization status cache;
[0140] A status change module, configured to change the status of the incremental synchronization service from paused to running based on the second indication information by the global synchronization status cache, so that the incremental synchronization service is restarted.
[0141] In an alternative embodiment, the data in the data lake message queue is collected based on the original service data.
[0142] Specifically, real-time incremental collection of the original service data is achieved through the Change Data Capture (CDC) technology.
[0143] It should be noted that when implementing its functions, the device provided in the above embodiment is only illustrated by the division of the above function modules. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0144] The data synchronization device of the data lake in the embodiment of the present application only pauses the inflow of incremental data corresponding to the first target data lake table during the data synchronization process, and the inflow of data in other data lake tables is not affected. Moreover, it can ensure the smooth connection of the stranded data synchronization and the incremental data synchronization, that is, it can improve the data synchronization efficiency on the basis of reducing the complexity of data management and the difficulty of data processing connection.
[0145] The embodiment of the present application also provides a data synchronization system for a data lake. Please refer to Figure 3, in the data synchronization system of this data lake, the upstream and downstream of data processing include: data collection, data synchronization, and data usage (i.e., the data warehouse system). During the data synchronization process, it is decoupled from the data collection end and the data usage end through a message middleware, and the message middleware uses Kafka (a high-throughput distributed publish-subscribe message system). The original business data is collected in real time in the form of CDC into the message queue connected to the data synchronization process, which is called the data lake message queue. During the data synchronization process, the collected data is obtained from the data lake message queue and synchronously to the tables in the data lake in real time, or further synchronously to the message queue connected to the data warehouse system, which is called the data warehouse message queue. The external data warehouse system obtains the synchronously data from the data warehouse message queue and performs its own business synchronization processing. That is, the data lake message queue collects the incremental data corresponding to the data lake tables from the original business data. When receiving a data synchronization request sent by the target data warehouse, the data lake message queue performs an initialization synchronization process when the data synchronization information indicates the initialization synchronization service information, and performs an incremental synchronization process when the data synchronization information indicates the incremental synchronization service information.
[0146] In this system, the processing of the data lake data synchronization process can specifically include: 3 services (incremental synchronization service, stock synchronization service, and stranded synchronization service), which are responsible for synchronizing incremental data, stock data, and stranded data respectively, 2 data transmission message queues (data lake message queue, data warehouse message queue), 2 task notification message queues (stock task message queue, stranded task message queue), and 1 global data synchronization status cache for caching the data synchronization status; among them, the 3 services can be implemented through programs or through relevant servers, and the specific description is as follows:
[0147] I. Incremental Synchronization Service
[0148] Responsible for obtaining the real-time collected incremental data from the data lake message queue and synchronizing it to the data lake and the data warehouse message queue. At the same time, as the main service in the entire data synchronization process, it coordinates the task execution of the stock synchronization service and the stranded synchronization service. The incremental synchronization service uses a polling method to pull data from the data lake message queue, but before each pull, it is necessary to ensure that in the global synchronization status cache, the status of the incremental synchronization task is in the "running" state, otherwise it is in the "paused" state, and no data pull will be performed.
[0149] II. Stock Synchronization Service
[0150] When the external data warehouse newly adds a synchronization requirement for the data lake, it is responsible for synchronizing the existing stock data in the data lake to the data warehouse message queue. During the synchronization process, the input of data collection for the synchronized data lake table is suspended.
[0151] III. Stranded Data Synchronization Service
[0152] Responsible for compensating and synchronizing the consumption of the data stranded in the data lake message queue due to the occurrence of inventory data synchronization after the inventory synchronization is completed.
[0153] IV. Global Synchronization Status Cache
[0154] It includes three parts:
[0155] (1) Synchronization relationship between the data lake table and the data warehouse object
[0156] Specifies that in addition to synchronizing the data of the collected data lake table to the corresponding table in the data lake, it also needs to be synchronized to the external data warehouse object.
[0157] (2) Inventory task status of the data lake table
[0158] Includes the number of inventory tasks currently contained in the data lake table, task_num, and the start and end offset positions of the data stranded in the message queue due to inventory synchronization: begin_offset, end_offset.
[0159] (3) Incremental synchronization task status
[0160] Specifies the start status of the current incremental synchronization task: running status, paused status. In the running status, the incremental synchronization service can consume the data in the data lake message queue. In the paused status, the incremental synchronization service cannot consume the data in the data lake message queue.
[0161] V. Data Lake Message Queue
[0162] Mainly serves as the data transmission channel for the data collection process and the data synchronization process. In addition, when the external data warehouse has a new data synchronization requirement for the data lake, it also serves as an event notification channel to notify the incremental synchronization service to update the synchronization relationship of the data lake table and the inventory synchronization status of the data lake table.
[0163] VI. Data Warehouse Message Queue
[0164] Serves as the data transmission channel between the data synchronization process and the external data warehouse.
[0165] VII. Inventory Task Message Queue
[0166] The channel through which the incremental synchronization service notifies the inventory synchronization service to start an inventory synchronization task.
[0167] VIII. Stranded Task Message Queue
[0168] The channel through which the incremental synchronization service notifies the stranded synchronization service to start a stranded synchronization task.
[0169] In the data synchronization process, according to the changes in the data synchronization requirements of the external data warehouse system for the data lake, data synchronization includes two scenarios: the incremental synchronization scenario in the stable state and the initialization synchronization scenario in the non-stable state. When the data synchronization requirements of the external data warehouse are stable and unchanged, data synchronization is in the incremental synchronization state, and the real-time collected data in the data lake message queue is consumed in real time by the incremental synchronization service and synchronized to the data lake and the data warehouse message queue. When the external data warehouse adds a data synchronization requirement for the data lake, the incremental synchronization state of the data involved in the data lake table is broken. To ensure the integrity and orderliness of the data, the data synchronization process for the external data warehouse object sequentially experiences the synchronization of the existing data (the existing data in the data lake), the synchronization of the stranded data (the data stranded in the data lake message queue), and finally returns to the synchronization of the incremental data (the incremental data in the data lake message queue), that is, during the access process, an initialization synchronization process is included. This system takes into account the above two synchronization scenarios, realizes the real-time incremental synchronization of the collected data, and at the same time, when the external data warehouse adds a data synchronization requirement, the processing volume of the collected data in the data lake message queue is not affected.
[0170] Please refer to Figure 4 , in the incremental synchronization process, when the data synchronization requirements of the external data warehouse system for the data lake do not change, and in the global synchronization status cache, the status of the incremental synchronization task is "running", the incremental synchronization task is executed. The incremental synchronization service pulls the original business data collected from the data lake message queue and checks in the global synchronization status cache whether there is an existing task being executed for the data lake table corresponding to the data, that is, the number of existing tasks is not 0. If so, it skips and does not synchronize the data, and this polling ends. Otherwise, it first synchronizes the data into the lake, and then goes to the global synchronization status cache to check whether there is a synchronization relationship between the data lake table and the external data warehouse object. If there is, it continues to synchronize to the data warehouse message queue.
[0171] Please refer to Figure 5 and Figure 6 , in the initialization synchronization process, when the external data warehouse adds a data synchronization requirement for the data lake table, a data synchronization request message is sent to the data lake message queue. The message contains the data lake table to be synchronized and the external data warehouse object to which it is synchronized. After receiving the message, the incremental synchronization service updates the synchronization relationship between the data lake table involved and the external data warehouse object in the global synchronization status cache, increments the number of existing synchronization tasks for the data lake table by 1, and if there was no existing synchronization task for this data lake table before, that is, the number of existing synchronization tasks is 0, it records the offset position of the current request message in the message queue as begin_offset. At the same time, it immediately notifies the existing synchronization service through the existing synchronization message queue to execute the synchronization of the existing data of the data lake table involved to the data warehouse message queue.
[0172] Since the number of stock synchronization tasks for the involved data lake table in the global synchronization status cache is not zero, when the incremental synchronization service subsequently pulls the collection data of the involved data lake table, no synchronization processing is performed.
[0173] For the stock synchronization service, when a stock synchronization task is completed, a message is immediately sent to the data lake message queue, indicating the data lake table for which the stock synchronization task is completed. For the incremental synchronization service, after receiving this message, in the global synchronization status cache, the number of stock synchronization tasks for this data lake table is decremented by 1. And after the decrement, if the current number of stock synchronization tasks for this data lake table happens to be zero, the offset position of the current message in the message queue is recorded as end_offset, and the status of the incremental synchronization task is set to "paused".
[0174] As mentioned above, if all the stock synchronization tasks for the current data lake table have been completed, it is necessary to restart the incremental ingestion of the collection data of this data lake table and synchronize it to the external data warehouse object. However, since during the stock synchronization process, the incremental synchronization service does not process the synchronization of the collection data of this data lake table, it can be considered that the data is detained in the data lake message queue. Therefore, if we want to restart the synchronization of incremental data, we must first synchronize the detained data of this data lake table. The position range of the detained data in the data lake message queue is determined by the begin_offset and end_offset recorded in the stock synchronization status of the data lake table in the global synchronization status cache.
[0175] The incremental synchronization service assembles the start and end offset positions of the above-mentioned detained data in the message queue and the data lake table name into a message and sends it to the detained synchronization message queue. The detained synchronization service, after receiving this message, performs targeted data synchronization consumption for the detained data of this data lake table in the data lake message queue according to the start and end offset positions provided in the message.
[0176] During the process of synchronizing the detained data, since in the global synchronization status cache, the status of the incremental synchronization task is set to "paused", before the task status is changed back to the "running" state, the task execution of the incremental synchronization service is always ignored and skipped during polling.
[0177] Finally, when the detained synchronization task is completed, the detained synchronization service resets the status of the incremental synchronization task in the global synchronization status cache to "running". The incremental synchronization service, after detecting that the task status is "running" during polling, restarts the incremental synchronization task.
[0178] Since the detained synchronization service also performs data synchronization consumption on the data lake message queue and only consumes the collection data of the previously detained data lake tables, for the data lake message queue, the data synchronization consumption is not paused and the data processing volume is not reduced.
[0179] An embodiment of the present application further provides an electronic device, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement the data synchronization method of the data lake provided in the above method embodiment.
[0180] The memory can be used to store software programs and modules. By running the software programs and modules stored in the memory, the processor can execute various functional applications and achieve high-level autonomous driving. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.
[0181] The method embodiment provided by the embodiment of the present application can be executed in a computer terminal, a server, or a similar computing device, that is, the above electronic device can include a computer terminal, a server, or a similar computing device. Figure 7 is a hardware structure block diagram of an electronic device that runs a data synchronization method of a data lake provided by an embodiment of the present application. As Figure 7 shown, the internal structure of the electronic device can include, but is not limited to: a processor, a network interface, and a memory. Among them, the processor, network interface, and memory in the electronic device can be connected through a bus or other means. In the embodiment of this specification Figure 7 it is taken as an example of being connected through a bus.
[0182] Among them, the processor (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device. The network interface may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.). The memory is the memory device in the electronic device, used to store programs and data. It can be understood that the memory here can be a high-speed RAM storage device, or a non-volatile memory device, such as at least one disk storage device; optionally, it can also be at least one storage device located far from the aforementioned processor. The memory provides a storage space, and this storage space stores the operating system of the electronic device, which may include but is not limited to: Windows system (an operating system), Linux (an operating system), Android (a mobile operating system) system, IOS (a mobile operating system) system, etc., and this application does not make any limitations in this regard; and, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space, and these instructions can be one or more computer programs (including program codes). In the embodiments of this specification, the processor loads and executes one or more instructions stored in the memory to implement the data synchronization method of the data lake provided in the above method embodiments.
[0183] The embodiments of this application also provide a computer-readable storage medium, in which at least one instruction or at least one segment of program is stored, and the at least one instruction or at least one segment of program is loaded and executed by the processor to implement the data synchronization method of the data lake provided in the method embodiments.
[0184] Optionally, in this embodiment, the above storage medium may include but is not limited to: USB flash drive, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disc and other various media that can store program codes.
[0185] It should be noted that: the above sequence of the embodiments of this application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired results. In certain embodiments, multi-small sample image classification and parallel processing are also possible or may be advantageous.
[0186] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments.
[0187] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0188] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data synchronization method for a data lake, characterized in that, Including: The external data warehouse sends a first data synchronization request to the data lake message queue; The first data synchronization relationship information is carried in the first data synchronization request; The incremental synchronization service obtains the first data synchronization request from the data lake message queue; Based on the first data synchronization request, the incremental synchronization service sends the first data synchronization relationship information to the global synchronization status cache to update the synchronization relationship in the global synchronization status cache and obtain a first target synchronization relationship; And based on the first data synchronization request, the incremental synchronization service sends a full synchronization request to the full synchronization service, so that the full synchronization service synchronizes the full data in the first target data lake table to the first target data warehouse corresponding to the first target data lake table; the full synchronization relationship information is carried in the full synchronization request, and the full synchronization relationship information is determined based on the first data synchronization relationship information; And based on the full synchronization request, the incremental synchronization service suspends the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue, so that the incremental data corresponding to the first target data lake table remains in the data lake message queue; When the incremental synchronization service queries that the synchronization progress information in the global synchronization status cache indicates that the full synchronization task corresponding to the first target data lake table has been completed, it sends a retention synchronization request to the retention synchronization service, so that the retention synchronization service sends the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the retention synchronization relationship information is carried in the retention synchronization request, and the retention synchronization relationship information is determined based on the first data synchronization relationship information; And suspend the incremental synchronization service; When the synchronization progress information indicates that the retention synchronization task corresponding to the first target data lake table has been completed, restart the incremental synchronization service, and the incremental synchronization service sends the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; Among them, the synchronization progress information is determined based on the full synchronization service, the retention synchronization service and the incremental synchronization service.
2. The data synchronization method for a data lake according to claim 1, characterized in that, The method includes two data synchronization scenarios, specifically including the incremental synchronization scenario in the steady state and the initialization synchronization scenario in the non-steady state. The steps from the external data warehouse sending the first data synchronization request to the data lake message queue to the incremental synchronization service sending the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizing it to the first target data warehouse are determined based on the initialization synchronization scenario in the non-steady state. The data synchronization steps in the incremental synchronization scenario in the steady state include: The external data warehouse sends a second data synchronization request to the data lake message queue; the second data synchronization relationship information is carried in the second data synchronization request; The incremental synchronization service obtains the second data synchronization request from the data lake message queue; Based on the second data synchronization request, the incremental synchronization service sends the second data synchronization relationship information to the global synchronization status cache to update the synchronization relationship in the global synchronization status cache, obtaining a second target synchronization relationship; Based on the second target synchronization relationship, the incremental synchronization service sends the incremental data corresponding to the newly added second target data lake table in the data lake message queue to the second target data lake table, or further synchronizes it to the second target data warehouse if there is a corresponding synchronization relationship between the second target data lake table and the second target data warehouse.
3. The data synchronization method for a data lake according to claim 1, characterized in that, The method further includes the step of determining the incremental data corresponding to the stranded first target data lake table. Determining the incremental data corresponding to the stranded first target data lake table includes: When the incremental synchronization service queries that the number of first inventory synchronization tasks corresponding to the first target data lake table in the synchronization progress information is zero, the incremental synchronization service obtains the first offset position of the incremental data corresponding to the first target data lake table in the data lake message queue and records the first offset position in the global synchronization status cache; the number of first inventory synchronization tasks is the number of inventory synchronization tasks corresponding to the first target data lake table already recorded in the global synchronization status cache before the incremental synchronization service obtains the first data synchronization request; When the incremental synchronization service queries that the number of second inventory synchronization tasks corresponding to the first target data lake table in the synchronization progress information is zero, the incremental synchronization service obtains the second offset position of the incremental data corresponding to the first target data lake table in the data lake message queue and records the second offset position in the global synchronization status cache; the number of second inventory synchronization tasks is the number of inventory synchronization tasks corresponding to the first target data lake table recorded in the global synchronization status cache after the incremental synchronization service obtains the first data synchronization request; The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache.
4. The data synchronization method for a data lake according to claim 3, characterized in that, The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the first offset position and the second offset position in the global synchronization status cache, including: The incremental synchronization service obtains the first offset position and the second offset position from the global synchronization status cache; The incremental synchronization service sends the stranded synchronization request to the stranded synchronization service; the stranded synchronization request also carries the first offset position and the second offset position; The stranded synchronization service determines the incremental data corresponding to the stranded first target data lake table based on the stranded synchronization request.
5. The data synchronization method for a data lake according to claim 3, characterized in that, The method further includes: When the number of the first stock synchronization tasks is not zero, the incremental synchronization service does not obtain the position of the first offset.
6. The data synchronization method for a data lake according to claim 3, characterized in that, The method further includes: When the number of the second stock synchronization tasks is not zero, the incremental synchronization service does not obtain the position of the second offset.
7. The data synchronization method for a data lake according to claim 1, characterized in that, The method further includes the step of determining that all the stock synchronization tasks corresponding to the first target data lake table have been completed. Determining that all the stock synchronization tasks corresponding to the first target data lake table have been completed includes: After completing one of the stock synchronization tasks, the stock synchronization service sends a first indication message to the data lake message queue; The data lake message queue receives the first indication message; The incremental synchronization service obtains the first indication message from the data lake message queue; The incremental synchronization service sends the first indication message to the global synchronization status cache; The global synchronization status cache updates the synchronization progress information based on the first indication message; Repeat the steps of completing one of the stock synchronization tasks until the global synchronization status cache updates the synchronization progress information based on the first indication message until the synchronization progress information indicates that the number of the stock synchronization tasks corresponding to the first target data lake table is zero, then it is determined that all the stock synchronization tasks corresponding to the first target data lake table have been completed.
8. A data synchronization device for a data lake, characterized in that, The device includes: A synchronization request sending module, configured to send a first data synchronization request to the data lake message queue by an external data warehouse; the first data synchronization request carries first data synchronization relationship information; A synchronization request obtaining module, configured to obtain the first data synchronization request from the data lake message queue by the incremental synchronization service; A stock synchronization module, configured to send, by the incremental synchronization service based on the first data synchronization request, the first data synchronization relationship information to the global synchronization status cache to update the synchronization relationship in the global synchronization status cache to obtain a first target synchronization relationship; and configured to send, by the incremental synchronization service based on the first data synchronization request, a stock synchronization request to the stock synchronization service to enable the stock synchronization service to synchronize the stock data in the first target data lake table to a first target data warehouse corresponding to the first target data lake table; the stock synchronization request carries stock synchronization relationship information, and the stock synchronization relationship information is determined based on the first data synchronization relationship information; and configured to pause, by the incremental synchronization service based on the stock synchronization request, the incremental synchronization of the incremental data corresponding to the first target data lake table in the data lake message queue to enable the incremental data corresponding to the first target data lake table to stay in the data lake message queue; A retention synchronization module, configured to send a retention synchronization request to a retention synchronization service when the synchronization progress information in the global synchronization status cache queried by the incremental synchronization service indicates that all the stock synchronization tasks corresponding to the first target data lake table have been completed, so that the retention synchronization service sends the incremental data corresponding to the first target data lake table retained in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse; the retention synchronization request carries retention synchronization relationship information, and the retention synchronization relationship information is determined based on the first data synchronization relationship information; and it is also configured to pause the incremental synchronization service. An incremental synchronization module, configured to restart the incremental synchronization service when the synchronization progress information indicates that all the retention synchronization tasks corresponding to the first target data lake table have been completed, and the incremental synchronization service sends the incremental data corresponding to the first target data lake table newly added in the data lake message queue to the first target data lake table and further synchronizes it to the first target data warehouse.
9. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the data synchronization method for a data lake as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing at least one instruction or at least one segment of a program, the at least one instruction or the at least one segment of the program being loaded and executed by a processor to implement the data synchronization method for a data lake according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data synchronization method and system
CN105183860A
Method and system for loading high-timeliness data into data lake
CN111367984A