Big data synchronization method and device, electronic equipment and storage medium

By creating external data tables and data loading channels on the distributed computing platform, the problems of slow synchronization speed and resource occupation are solved, and efficient data synchronization is achieved.

CN120407681APending Publication Date: 2025-08-01JIAXING JUSHUITAN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510505199.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, the distributed computing platform synchronizes data to the target database with a slower speed and low efficiency, and occupies more resources.

Method used

Create external data tables in the target distributed computing platform, write data to the target cloud storage service, and establish communication connections with the target distributed database and cloud storage service through script files, create a data loading channel, and synchronize the data to the target distributed database.

Benefits of technology

It realizes the rapid synchronization of large amounts of data to the target database, improving synchronization efficiency without occupying too many resources on the distributed computing platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407681A_ABST
    Figure CN120407681A_ABST
Patent Text Reader

Abstract

The invention discloses a big data synchronization method and device, electronic equipment and a storage medium, and relates to the technical field of data processing. The method comprises the following steps: creating a first external data table on a target distributed computing platform, and writing to-be-synchronized data in the target distributed computing platform into a target cloud storage service associated with the first external data table; establishing a communication connection with a target distributed database based on the first script file, and establishing a communication connection with the target cloud storage service based on the second script file; and creating a target data loading channel, and synchronizing the to-be-synchronized data stored in the target cloud storage service to a target distributed database based on the target data loading channel. According to the scheme, a large amount of data in the distributed computing platform can be quickly synchronized to the target database, and resources of the distributed computing platform cannot be excessively occupied while the synchronization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, electronic device and storage medium for big data synchronization. Background Art

[0002] With the rapid development of big data technology, the demand of enterprises for cross-platform and cross-system data synchronization is increasing day by day. At present, the data in the distributed computing platform is mainly synchronized to the target database directly through database technology. However, this method has a slow synchronization speed, low efficiency and occupies a large amount of resources of the distributed computing platform.

[0003] How to quickly synchronize a large amount of data in the distributed computing platform to the target database, improve the synchronization efficiency and not occupy too much resources of the distributed computing platform is the key issue studied in the industry. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and storage medium for big data synchronization, so as to quickly synchronize a large amount of data in the distributed computing platform to the target database, improve the synchronization efficiency and not occupy too much resources of the distributed computing platform.

[0005] According to one aspect of the present invention, there is provided a method for big data synchronization, the method comprising:

[0006] Create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table;

[0007] Create a communication connection with the target distributed database based on the first script file, and create a communication connection with the target cloud storage service based on the second script file;

[0008] Create a target data loading channel, and synchronize the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel.

[0009] According to another aspect of the present invention, there is provided a device for big data synchronization, the device comprising:

[0010] A first external data table creation module, configured to create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table;

[0011] A communication connection creation module, configured to create a communication connection with the target distributed database based on the first script file, and create a communication connection with the target cloud storage service based on the second script file;

[0012] A data loading channel creation module, configured to create a target data loading channel, and synchronize the data to be synchronized stored in the target cloud storage service to a target distributed database based on the target data loading channel.

[0013] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the big data synchronization method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the big data synchronization method according to any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, there is provided a computer program product including a computer program, which implements the big data synchronization method according to any embodiment of the present invention when executed by a processor.

[0019] The technical solution of the embodiment of the present invention creates a first external data table in a target distributed computing platform, writes the data to be synchronized in the target distributed computing platform into a target cloud storage service associated with the first external data table; creates a communication connection with the target distributed database based on a first script file, and creates a communication connection with the target cloud storage service based on a second script file; creates a target data loading channel, and synchronizes the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel, so that a large amount of data in the distributed computing platform can be quickly synchronized to the target database, improving the synchronization efficiency and not occupying too much resources of the distributed computing platform at the same time.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0022] Figure 1 is a flowchart of a big data synchronization method provided in Embodiment 1 of the present invention;

[0023] Figure 2 is a flowchart of a big data synchronization method provided in Embodiment 2 of the present invention;

[0024] Figure 3 is a flowchart of a big data synchronization method provided in Embodiment 3 of the present invention;

[0025] Figure 4 is a flowchart of another big data synchronization method provided in Embodiment 3 of the present invention;

[0026] Figure 5 is a schematic diagram of the specific steps of a big data synchronization solution provided in Embodiment 3 of the present invention;

[0027] Figure 6 is a schematic diagram of the structure of a big data synchronization device provided in Embodiment 4 of the present invention;

[0028] Figure 7 is a schematic diagram of the structure of an electronic device for implementing the big data synchronization method of the embodiments of the present invention. Detailed Embodiments

[0029] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Embodiment 1

[0032] Figure 1 is a flowchart of a big data synchronization method provided according to Embodiment 1 of the present invention. This embodiment is applicable to the situation of synchronizing a large amount of data in a distributed computing platform to a target database for storage. This method can be executed by a big data synchronization device, which can be implemented in the form of hardware and / or software, and the big data synchronization device can be configured in an electronic device such as a computer, a server or a tablet computer. As Figure 1 shown, the method includes:

[0033] Step 110, create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into a target cloud storage service associated with the first external data table.

[0034] Among them, the target distributed computing platform can be any software system deployed locally or in the cloud that can be used to execute distributed computing tasks, and it is not limited in this embodiment.

[0035] In this embodiment, the first external data table can be an external data table of an Object Storage Service (OSS), or an external data table of other storage services, and it is not limited in this embodiment.

[0036] It should be noted that the data stored in the first external data table created in this embodiment is not the data to be synchronized itself, but the definition or mapping of metadata; for example, if the first external data table is an OSS external data table, then it can contain information on how to access the data in the OSS service in the cloud, such as the path, format (for example, Parquet, ORC or CSV, etc.) of the data file, column definition, and possible data partition information.

[0037] Optionally, in this embodiment, when there is a large amount of business data (such as order data, freight calculation results, or profit analysis results, etc.) in the target distributed computing platform that needs to be synchronized to the target distributed database for storage, a first external data table can be created in the target distributed computing platform; further, the data to be synchronized in the target distributed computing platform can be written into the target cloud storage service associated with the first external data table based on the first external data table; for example, in the above example, the data to be synchronized can be written into the corresponding path of the OSS service based on the path of the data file stored in the OSS external data table.

[0038] Among them, the data to be synchronized can be real-time data streams, structured data, unstructured data, or log data, etc., which are not limited in this embodiment; it can be understood that in this embodiment, the data to be synchronized can be big data with a large amount of data, diverse data types, fast update speed, and relatively high data value density.

[0039] Step 120: Create a communication connection with the target distributed database based on the first script file, and create a communication connection with the target cloud storage service based on the second script file.

[0040] Optionally, in this embodiment, after writing the data to be synchronized into the target cloud storage service, a communication connection with the target distributed database can be further created based on the first script file, and a communication connection with the target cloud storage service can be created based on the second script file; among them, the first script file and the second script file can be python script files, or script files written in other programming languages, which are not limited in this embodiment; it should be noted that in this embodiment, the first script file and the second script file can also be the same script file. In this embodiment, only for the convenience of narration, the different functions of the script file are described as different script files, which is not a limitation of this embodiment.

[0041] Optionally, in this embodiment, creating a communication connection with the target distributed database based on the first script file and creating a communication connection with the target cloud storage service based on the second script file can include: obtaining the attribute information of the target distributed database based on the first function of the first script file; the attribute information includes at least one of the following: server address, port, username, user password, and the name of the target distributed database; returning a connection object connecting to the target distributed database; determining whether the data to be synchronized exists in the target cloud storage service based on the second function of the second script file.

[0042] Among them, the first function of the first script file is the function that can obtain the attribute information of the target distributed database. For example, it can obtain information such as the server address, port, username, user password, and the name of the target distributed database. Further, it can return a connection object for connecting to the target distributed database.

[0043] Exemplarily, in this embodiment, the process of creating a communication connection with the target cloud storage service can be implemented through the following script file. The "def_get_conn" in this script file is the first function involved in this embodiment:

[0044] def_get_conn(self,host,port,user,password,database):

[0045] conn=pymysql.connect(host=host,port=port,user=user,password=password,database=database,charset='utf8mb4').

[0046] In this embodiment, the second function of the second script file is the function that can determine whether the to-be-synchronized data exists in the target cloud storage service; Exemplarily, in this embodiment, the process of creating a communication connection with the target cloud storage service can be implemented through the following script file. The "defoss_exists" in this script file is the second function involved in this embodiment:

[0047]

[0048]

[0049] Step 130, create a target data loading channel, and synchronize the to-be-synchronized data stored in the target cloud storage service to the target distributed database based on the target data loading channel.

[0050] Optionally, in this embodiment, after creating a communication connection with the target distributed database based on the first script file and creating a communication connection with the target cloud storage service based on the second script file, a target data loading channel can be further created, and further, the to-be-synchronized data stored in the target cloud storage service can be synchronized to the target distributed database based on the target data loading channel.

[0051] It can be understood that the target data loading channel is the data channel between the target cloud storage service and the target distributed database. When synchronizing various data in the distributed computing platform, it does not need to occupy the computing resources of the target distributed computing platform and will not affect the computing efficiency of the target distributed computing platform.

[0052] Optionally, in this embodiment, before synchronizing the data to be synchronized based on the target data loading channel, the number of partitions of the target distributed database can also be determined based on the data volume of the data to be synchronized, and the data to be synchronized is synchronized to different partitions of the target distributed database based on the target data loading channel; the number of partitions of the target distributed database can also be determined based on the generation time of each data in the data to be synchronized, and the data to be synchronized generated on the same day is synchronized to the same partition of the target distributed database based on the target data loading channel.

[0053] The technical solution of this embodiment creates a first external data table in the target distributed computing platform, writes the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table; creates a communication connection with the target distributed database based on the first script file, and creates a communication connection with the target cloud storage service based on the second script file; creates a target data loading channel, and synchronizes the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel, which can quickly synchronize a large amount of data in the distributed computing platform to the target database, improving the synchronization efficiency without occupying too many resources of the distributed computing platform.

[0054] Embodiment 2

[0055] Figure 2 FIG. is a flowchart of a big data synchronization method according to Embodiment 2 of the present invention. This embodiment further refines the above technical solution, and the technical solution in this embodiment can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the method includes:

[0056] Step 210, create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table.

[0057] Optionally, in this embodiment, creating a first external data table in the target distributed computing platform and writing the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table may include: determining a target data storage format and a target storage path; the target storage path is the storage path in the target cloud storage service; creating the first external data table based on the target data storage format and the target storage path; obtaining the target storage path, and writing the data to be synchronized into the target storage path of the target cloud storage service by means of insertion.

[0058] Among them, the target data storage format may be one or more of a text format, a serialization format, a binary format, a log format, or other formats, which are not limited in this embodiment; the target storage path may be the storage path in the target cloud storage service, that is, the actual storage path of the data to be synchronized.

[0059] In an alternative implementation of this embodiment, after determining the target data storage format of the data to be synchronized and the target storage path in the target cloud storage service, a first external data table may be further created based on the target data storage format and the target storage path, where the first external data table may be an OSS external data table.

[0060] Furthermore, the target storage path recorded in the first external data table may be obtained, and the data to be synchronized is written into the target storage path of the target cloud storage service by means of insertion; it should be noted that the number of target storage paths in this embodiment may also be one or more, which are not limited in this embodiment.

[0061] Step 220: Create a communication connection with the target distributed database based on the first script file, and create a communication connection with the target cloud storage service based on the second script file.

[0062] Step 230: Construct a data loading statement based on the third script file, and determine the target parameters; execute the data loading statement to load the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database.

[0063] Among them, the target parameters include at least one of the following: data source path, target table name, and field mapping relationship; the target partition is determined by the target table name and / or the field mapping relationship.

[0064] Optionally, in this embodiment, after creating a communication connection with the target distributed database based on the first script file and creating a communication connection with the target cloud storage service based on the second script file, a data loading statement can be further constructed based on the third script file to determine the data source path, the target table name, or the field mapping relationship, and execute the data loading statement, so as to load the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database.

[0065] Among them, the third script file can be a Python script file or a script file written in other programming languages, which is not limited in this embodiment; it should be noted that the third script file in this embodiment can be the same script file as the first script file and the second script file. In this embodiment, only for the convenience of narration, the different functions of the script file are described as different script files, which is not a limitation of this embodiment.

[0066] Exemplarily, the third script file is implemented by the following code, where the data loading statement can be def_broker_load;

[0067]

[0068]

[0069]

[0070] Correspondingly, in this embodiment, before loading the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database, it may further include: updating the partition of the target distributed database based on the data volume and / or data generation time of the data to be synchronized; determining the target partition based on the partition result; the partition result includes the number of partitions or the attribute information of the partition; among them, the attribute information of the partition may include the location information of the partition in the target distributed database, or may include the generation time of the stored data, etc., which is not limited in this embodiment.

[0071] Optionally, in this embodiment, after constructing the data loading statement based on the third script file, that is, creating the target data loading channel, the partition of the target distributed database can be further updated based on the data volume, generation time, or data volume and generation time of the data to be synchronized; further, the target partition can be determined based on the obtained partition result; the partition result includes the number of partitions or the attribute information of the partition.

[0072] Step 240: Obtain the task tags generated during the execution process of the data loading statement in real time, and determine the task status based on the task tags; in the case where the task status is task cancellation, send an exception message to the target user and give an alarm.

[0073] Among them, the task status includes: task completion, task cancellation, or waiting for execution.

[0074] Optionally, in this embodiment, during the process of loading the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database, the execution process of the data loading statement can be monitored, that is, the task tags generated during the execution process of the data loading statement can be obtained in real time, and the task status can be determined based on the task tags; for example, if the obtained task tag is "finished", then the current task status is task completion; if the obtained task tag is "cancelled", then the current task status is task cancellation; if the obtained task tag is "wait for 5 seconds", then the current task status is waiting for execution.

[0075] Furthermore, if it is determined that the task status is task cancellation and the error message is not 'type:LOAD_RUN_FAIL; msg:all partitions have no load data' or 'type:LOAD_RUN_FAIL; msg:No partitions have data available for loading', it can be considered an abnormal situation, and at this time, an alarm can be triggered. Specifically, an exception message can be sent to the target user and an alarm can be given. For example, an alarm message can be sent to the user through the target application or applet. The target application can be any authorized interactive application installed on the mobile terminal, and it is not limited in this embodiment.

[0076] In another alternative implementation of this embodiment, if it is determined that the task status is waiting for execution, then the current retry count can be further determined. If the data synchronization task is still not completed when the current retry count is greater than the set retry count threshold, an alarm can be triggered.

[0077] In another alternative implementation of this embodiment, if it is determined that the task status is waiting for execution, the waiting duration can be further determined. If the data synchronization task is still not completed when the waiting duration is greater than or equal to the set waiting duration, an alarm can be triggered.

[0078] The solution of this embodiment can construct a data loading statement based on a third script file to determine target parameters; execute the data loading statement to load the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database. The data to be synchronized can be stored in the template partition of the target distributed database by constructing a corresponding data channel. Meanwhile, the task tag generated during the execution process of the data loading statement is obtained in real time, and the task status is determined based on the task tag; in the case where the task status is task cancellation, an exception message is sent to the target user and an alarm is given, so as to monitor the data synchronization task, respond to different task statuses in real time, and provide a basis for improving the synchronization efficiency of the data synchronization task.

[0079] Embodiment III

[0080] Figure 3 FIG. is a flowchart of a big data synchronization method provided according to Embodiment III of the present invention. This embodiment further refines the above technical solution, and the technical solution in this embodiment can be combined with each optional solution in one or more of the above embodiments. As Figure 3 shown, the method includes:

[0081] Step 310: Create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table.

[0082] Step 320: Obtain the attribute information of the target distributed database based on the first function of the first script file; return a connection object for connecting to the target distributed database; determine whether the data to be synchronized exists in the target cloud storage service based on the second function of the second script file.

[0083] Wherein, the attribute information includes at least one of the following: server address, port, user name, user password, and the name of the target distributed database.

[0084] Optionally, in this embodiment, the server address, port, user name, user password, or the name of the target distributed database and other attribute information of the target distributed database can be obtained based on the first function in the first script file, so as to return a connection object for connecting to the target distributed database; further, it can be determined whether the data to be synchronized exists in the target cloud storage service based on the second function of the second script file, that is, to determine the consistency between the data mapped by the first external data table and the data stored in the target cloud storage service.

[0085] Step 330: Determine the consistency between the data to be synchronized and the data stored in the target cloud storage service; in the case where it is determined that the data to be synchronized is inconsistent with the data stored in the target cloud storage service, determine the candidate data to be synchronized, create a target data loading channel, and synchronize the candidate data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel; in the case where it is determined that the data to be synchronized is consistent with the data stored in the target cloud storage service, stop the synchronization task of the data to be synchronized.

[0086] Optionally, in this embodiment, in the case where it is determined that the data mapped by the first external data table is inconsistent with the data stored in the target cloud storage service, a target data loading channel can be directly created, and all the data to be synchronized stored in the target cloud storage service can be synchronized to the target distributed database based on the target data loading channel.

[0087] In another alternative implementation of this embodiment, in the case where it is determined that the data to be synchronized is consistent with the data stored in the target cloud storage service, stop the synchronization task of the data to be synchronized, wait for the newly added synchronization data, and synchronize the newly added synchronization data.

[0088] The solution of this embodiment can quickly determine the data that needs to be synchronized through the data loading channel by verifying the consistency between the data mapped in the first external data table and the data stored in the target cloud storage service, improving the timeliness of data synchronization.

[0089] Step 340: In response to an instruction to increase the data to be synchronized, obtain the newly added data to be synchronized; determine the reference partition that matches the newly added data to be synchronized based on the data volume and / or data generation time of the newly added data to be synchronized; synchronize the newly added data to be synchronized to the reference partition of the target distributed database based on the target data loading channel.

[0090] Optionally, in this embodiment, after receiving the instruction to increase the data to be synchronized, the newly added data to be synchronized can be further obtained. It can be understood that the newly added data to be synchronized can be the intermediate data or result data generated by the target distributed computing platform continuously performing data calculation tasks.

[0091] Furthermore, the reference partition in the target distributed database that matches each newly added data to be synchronized can be determined based on the generation time of the newly added data to be synchronized, or the service identifier that generates the newly added data to be synchronized, and the newly added data to be synchronized can be synchronized to the reference partition of the target distributed database based on the pre-created target data loading channel.

[0092] The solution of this embodiment can achieve the specified storage of newly added data to be synchronized, and can realize the data synchronization between the target distributed database and the target distributed computing platform by obtaining the newly added data to be synchronized, determining the reference partition that matches the newly added data to be synchronized based on the data volume and / or data generation time of the newly added data to be synchronized, and synchronizing the newly added data to be synchronized into the reference partition of the target distributed database through the target data loading channel.

[0093] To better understand the big data synchronization method involved in this embodiment, Figure 4 is a flowchart of another big data synchronization method provided according to Embodiment III of the present invention; as Figure 4 shown, it mainly includes data preparation, data channel, data synchronization, and exception capture processes;

[0094] Among them, data preparation is mainly responsible for writing the data of the target distributed computing platform into the target cloud storage service; the data channel is responsible for constructing the connection between the target distributed database and the target cloud storage service; data loading is responsible for loading the data in the target cloud storage service into the target distributed database through the target data loading channel; exception monitoring mainly monitors the task status of each stage and alarms the abnormal tasks.

[0095] Figure 5 is a schematic diagram of the specific steps of a big data synchronization solution provided according to Embodiment III of the present invention; it mainly includes table mapping, updating the number of partitions, and writing to the corresponding partitions; among them, table mapping is mainly the mapping between the created first external data table and the target cloud storage service; updating the number of partitions refers to the number of partitions that need to be synchronized and updated this time. Further, different data can be written into different partitions.

[0096] The solution of this embodiment synchronizes the data in the distributed computing platform to the distributed database through the data loading channel, with a relatively fast synchronization speed and without occupying too many resources of the distributed computing platform. By creating an external table, it reduces the data storage of the distributed computing platform, and at the same time reduces the resource consumption of the data integration resource group and optimizes the cost.

[0097] Embodiment IV

[0098] Figure 6 is a schematic structural diagram of a big data synchronization device provided according to Embodiment IV of the present invention. As Figure 6 shown, the device includes: a first external data table creation module 610, a communication connection creation module 620, and a data loading channel creation module 630.

[0099] Among them, the first external data table creation module 610 is used to create a first external data table in the target distributed computing platform and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table;

[0100] The communication connection creation module 620 is used to create a communication connection with the target distributed database based on the first script file and create a communication connection with the target cloud storage service based on the second script file;

[0101] The data loading channel creation module 630 is used to create a target data loading channel and synchronize the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel.

[0102] In the solution of this embodiment, the first external data table creation module creates a first external data table in the target distributed computing platform and writes the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table; the communication connection creation module is used to create a communication connection with the target distributed database based on the first script file and create a communication connection with the target cloud storage service based on the second script file; the data loading channel creation module is used to create a target data loading channel and synchronize the data to be synchronized stored in the target cloud storage service to the target distributed database, which can quickly synchronize a large amount of data in the distributed computing platform to the target database, improve the synchronization efficiency and will not occupy too much resources of the distributed computing platform.

[0103] In an optional implementation manner of this embodiment, the first external data table creation module 610 is specifically used to determine the target data storage format and the target storage path; the target storage path is the storage path in the target cloud storage service;

[0104] Create the first external data table based on the target data storage format and the target storage path;

[0105] Obtain the target storage path and write the data to be synchronized to the target storage path of the target cloud storage service in an insert manner.

[0106] In an optional implementation manner of this embodiment, the communication connection creation module 620 is specifically used to obtain the attribute information of the target distributed database based on the first function of the first script file; the attribute information includes at least one of the following: server address, port, username, user password, and the name of the target distributed database;

[0107] Return a connection object for connecting to the target distributed database;

[0108] Determine whether the to-be-synchronized data exists in the target cloud storage service based on the second function of the second script file.

[0109] In an alternative implementation of this embodiment, the data loading channel creation module 630 is specifically configured to determine the consistency between the to-be-synchronized data and the data stored in the target cloud storage service.

[0110] In the case where it is determined that the to-be-synchronized data is inconsistent with the data stored in the target cloud storage service, determine candidate to-be-synchronized data, create a target data loading channel, and synchronize the candidate to-be-synchronized data stored in the target cloud storage service to the target distributed database based on the target data loading channel.

[0111] In the case where it is determined that the to-be-synchronized data is consistent with the data stored in the target cloud storage service, stop the synchronization task of the to-be-synchronized data.

[0112] In an alternative implementation of this embodiment, the data loading channel creation module 630 is further specifically configured to construct a data loading statement based on a third script file and determine target parameters.

[0113] Wherein, the target parameters include at least one of the following: data source path, target table name, and field mapping relationship.

[0114] Execute the data loading statement to load the to-be-synchronized data stored in the target cloud storage service into the target partition of the target distributed database; the target partition is determined by the target table name and / or the field mapping relationship.

[0115] Correspondingly, the big data synchronization device further includes: a partitioning unit, configured to update the partition of the target distributed database based on the data volume and / or data generation time of the to-be-synchronized data.

[0116] Determine the target partition based on the partitioning result; the partitioning result includes the number of partitions or the attribute information of the partitions.

[0117] In an alternative implementation of this embodiment, the big data synchronization device further includes: an alarm module, configured to obtain in real time the task tags generated during the execution process of the data loading statement, and determine the task status based on the task tags.

[0118] Wherein, the task status includes: task completed, task cancelled, or waiting to be executed.

[0119] In the case where the task status is task cancelled, send an exception message to the target user and give an alarm.

[0120] In an alternative implementation of this embodiment, the big data synchronization device further includes: a newly added data to be synchronized synchronization module, configured to obtain newly added data to be synchronized in response to an instruction to add data to be synchronized;

[0121] Determine a reference partition that matches the newly added data to be synchronized based on the data volume and / or data generation time of the newly added data to be synchronized;

[0122] Synchronize the newly added data to be synchronized to the reference partition of the target distributed database based on the target data loading channel.

[0123] The big data synchronization device provided by the embodiments of the present invention can execute the big data synchronization method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0124] In the technical solution of the embodiments of the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of the data to be synchronized in the distributed computing platform all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0125] Embodiment Five

[0126] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0127] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0128] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0129] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the big data synchronization method, which includes: creating a first external data table in a target distributed computing platform and writing the data to be synchronized in the target distributed computing platform into a target cloud storage service associated with the first external data table; creating a communication connection with a target distributed database based on a first script file and creating a communication connection with the target cloud storage service based on a second script file; creating a target data loading channel and synchronizing the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel.

[0130] In some embodiments, the big data synchronization method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the big data synchronization method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the big data synchronization method by any other suitable means (e.g., by means of firmware).

[0131] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0132] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0133] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0136] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0137] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.

[0138] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

[0139] An embodiment of the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the method for detecting a database provided in any embodiment of the present application.

[0140] In the process of implementing the computer program product, computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0141] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0142] Note that the above is only a preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A big data synchronization method, characterized in that, Including: Create a first external data table on the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table; Create a communication connection with the target distributed database based on the first script file, and create a communication connection with the target cloud storage service based on the second script file; Create a target data loading channel, and synchronize the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel.

2. The big data synchronization method according to claim 1, wherein The step of creating a first external data table on the target distributed computing platform and writing the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table includes: Determine the target data storage format and the target storage path; the target storage path is the storage path in the target cloud storage service; Create the first external data table based on the target data storage format and the target storage path; Obtain the target storage path, and write the data to be synchronized into the target storage path of the target cloud storage service by means of insertion.

3. The big data synchronization method according to claim 1, characterized in that, The step of creating a communication connection with the target distributed database based on the first script file and creating a communication connection with the target cloud storage service based on the second script file includes: Obtain the attribute information of the target distributed database based on the first function of the first script file; the attribute information includes at least one of the following: server address, port, user name, user password, and the name of the target distributed database; Return a connection object for connecting to the target distributed database; Determine whether the data to be synchronized exists in the target cloud storage service based on the second function of the second script file.

4. The big data synchronization method according to claim 3, wherein The step of creating a target data loading channel and synchronizing the data to be synchronized stored in the target cloud storage service to the target distributed database includes: Determine the consistency between the data to be synchronized and the data stored in the target cloud storage service; In the case where it is determined that the data to be synchronized is inconsistent with the data stored in the target cloud storage service, determine candidate data to be synchronized, create a target data loading channel, and synchronize the candidate data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel; In the case where it is determined that the data to be synchronized is consistent with the data stored in the target cloud storage service, stop the synchronization task of the data to be synchronized.

5. The big data synchronization method according to claim 1, wherein The step of creating a target data loading channel and synchronizing the data to be synchronized stored in the target cloud storage service to the target distributed database includes: Construct a data loading statement based on the third script file and determine the target parameters; Wherein, the target parameters include at least one of the following: data source path, target table name, and field mapping relationship; Execute the data loading statement, and load the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database; the target partition is determined by the target table name and / or the field mapping relationship. Correspondingly, before loading the data to be synchronized stored in the target cloud storage service into the target partition of the target distributed database, it further includes: Updating the partition of the target distributed database based on the data volume and / or data generation time of the data to be synchronized; Determining the target partition based on the partitioning result; the partitioning result includes the number of partitions or the attribute information of the partitions.

6. The big data synchronization method according to claim 5, wherein The method further includes: Obtaining in real time the task tags generated during the execution process of the data loading statement, and determining the task status based on the task tags; Wherein, the task status includes: task completed, task cancelled, or waiting to be executed; In the case where the task status is task cancelled, sending an exception message to the target user and giving an alarm.

7. The big data synchronization method according to claim 5, characterized in that, The method further includes: Responding to an instruction for increasing the data to be synchronized, and obtaining the newly added data to be synchronized; Determining a reference partition matching the newly added data to be synchronized based on the data volume and / or data generation time of the newly added data to be synchronized; Synchronizing the newly added data to be synchronized to the reference partition of the target distributed database based on the target data loading channel.

8. A big data synchronization device, characterized in that, It includes: A first external data table creation module, configured to create a first external data table in the target distributed computing platform, and write the data to be synchronized in the target distributed computing platform into the target cloud storage service associated with the first external data table; A communication connection creation module, configured to create a communication connection with the target distributed database based on a first script file, and create a communication connection with the target cloud storage service based on a second script file; A data loading channel creation module, configured to create a target data loading channel, and synchronize the data to be synchronized stored in the target cloud storage service to the target distributed database based on the target data loading channel.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the big data synchronization method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the big data synchronization method according to any one of claims 1-7 when executed.