Data synchronization method and device based on non-full log completion, equipment and medium

By building configuration files, obtaining corresponding components, and using distributors and fill operators to automatically fill part-time logs and merge data, solving the problem that part-time logs cannot be automatically filled during data synchronization, and achieving efficient and accurate data synchronization.

CN120196681APending Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344408.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

During the data synchronization process, the part-time log cannot be automatically filled, resulting in data loss or inaccuracy during data synchronization.

Method used

By building configuration files, the consumer, the distributor, the complement operator and the output device are obtained. The distributor is used to detect the log type and write it to the array. The complement operator is started asynchronously to complete the original array, and merge the data through the output device and then sent it to the target database.

Benefits of technology

It realizes automatic filling of part-time logs during data synchronization, ensures the accuracy of data synchronization, and avoids data blocking through concurrent operations, improving data writing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196681A_ABST
    Figure CN120196681A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data and finance, and provides a data synchronization method, device, equipment and medium based on non-full log completion, on one hand, a configuration file is constructed and published to obtain a consumer, a distributor, a completion operator and an output device, non-coding is realized, and the labor cost is effectively reduced; on one hand, the dispatcher asynchronously starts the completion operator, the completion operator is used for completing each piece of data in an original array, automatic completion of non-full logs can be achieved during data synchronization, and therefore the accuracy of data synchronization is guaranteed, and in addition, the concurrent operation can also achieve rapid completion of log data, and data blocking is avoided; and on the other hand, the output device is used for merging the data in the complementary array and then issuing the merged data to the target database, so that unnecessary data can be prevented from being issued to the target database, and the data writing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data and financial technology, and in particular to a data synchronization method, device, equipment and medium based on non-full log completion. Background Art

[0002] Data is the oil of the new era. Real-time analysis and computing play a huge role in today's society, especially in financial fields such as e-commerce, smart cities, and bank risk control. Mainstream real-time data analysis is not performed directly in the business database (i.e., OLTP (On-Line Transaction Processing)), which will put pressure on the business database. At the same time, the business database generally only has primary analysis capabilities. If the amount of data is large (such as TB (Terabyte) level), the single business database, whether CPU (Central Processing Unit) or memory, cannot support it. Therefore, a real-time big data synchronization tool is needed to synchronize data in real time (such as seconds-level delay) to the analytical database OLAP (Online Analytical Processing).

[0003] The principle of real-time data synchronization is mainly to collect the logs (binlog) generated by OLTP (On-Line Transaction Processing), organize the logs into the format required by the complete OLAP table, and then write them. However, the OLTP system is generally a business core library. In order to improve performance, it will not write the complete log to the binlog, but only write the values ​​of the updated part. This kind of log is called "non-full log". For example: a table has three fields id, name, and age. If only the age field value is updated at a certain time, then only the id and age information will appear in the log.

[0004] In order to improve the system throughput and reading speed, most OLAP engines write in batches when processing data intake. When processing incomplete data, they will automatically fill it with null values ​​(that is, invalid values). For example, in the example above, if you synchronize by conventional means, the name field will be set to null.

[0005] Therefore, completing incomplete logs during data synchronization has become an urgent problem to be solved. Summary of the invention

[0006] In view of the above, it is necessary to provide a data synchronization method, device, equipment and medium based on incomplete log completion, aiming to solve the problem that incomplete logs cannot be automatically completed during data synchronization.

[0007] A data synchronization method based on non-full log completion, the data synchronization method based on non-full log completion includes:

[0008] In response to a data synchronization instruction from a source database to a target database, construct and publish a configuration file to obtain a consumer, a distributor, a completion operator, and an outputter;

[0009] Obtain the table structure information of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain the target table structure of the target database;

[0010] Use the consumer to consume the log data in the source database;

[0011] Use the distributor to detect the log type of the log data, and write the log data into a pre-constructed original array or completion array according to the log type;

[0012] Use the distributor to asynchronously start the completion operator, and use the completion operator to perform completion processing on each piece of data in the original array, and write each piece of data obtained after the completion processing into the completion array;

[0013] Start the outputter according to a start strategy, and use the outputter to merge the data in the completion array to obtain target data, and send the target data to the target database.

[0014] A data synchronization device based on non-full log completion, the data synchronization device based on non-full log completion includes:

[0015] A publishing unit, configured to construct and publish a configuration file in response to a data synchronization instruction from a source database to a target database to obtain a consumer, a distributor, a completion operator, and an outputter;

[0016] An initialization unit, configured to obtain the table structure information of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain the target table structure of the target database;

[0017] A consumption unit, configured to use the consumer to consume the log data in the source database;

[0018] A writing unit, configured to use the distributor to detect the log type of the log data, and write the log data into a pre-constructed original array or completion array according to the log type;

[0019] A completion unit, configured to asynchronously start the completion operator by using the distributor, perform completion processing on each piece of data in the original array by using the completion operator, and write each piece of data obtained after the completion processing into the completed array;

[0020] A merging unit, configured to start the outputter according to a start policy, merge the data in the completed array by using the outputter to obtain target data, and send the target data to the target database.

[0021] A computer device, the computer device includes:

[0022] A memory, storing at least one instruction; and

[0023] A processor, executing the instruction stored in the memory to implement the data synchronization method based on non-full log completion.

[0024] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the data synchronization method based on non-full log completion.

[0025] It can be seen from the above technical solutions that, on the one hand, constructing and publishing a configuration file to obtain a consumer, a distributor, a completion operator, and an outputter realizes code-free, effectively reducing the labor cost; on the one hand, the distributor asynchronously starts the completion operator, and uses the completion operator to perform completion processing on each piece of data in the original array, which can automatically complete non-full logs during data synchronization, thus ensuring the accuracy of data synchronization. Moreover, concurrent operations can also achieve fast completion of log data and avoid data blocking; on the other hand, using the outputter to merge the data in the completed array and then send it to the target database can avoid sending unnecessary data to the target database and improve the data writing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of a preferred embodiment of the data synchronization method based on non-full log completion of the present invention.

[0027] Figure 2 is a functional module diagram of a preferred embodiment of the data synchronization device based on non-full log completion of the present invention.

[0028] Figure 3 is a schematic structural diagram of a computer device of a preferred embodiment of the data synchronization method based on non-full log completion of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] As Figure 1 shown, it is a flowchart of a preferred embodiment of the data synchronization method based on non-full log completion according to the present invention. According to different requirements, the order of steps in this flowchart can be changed and some steps can be omitted.

[0031] The data synchronization method based on non-full log completion is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0032] The computer device can be any electronic product that can perform human-computer interaction with users. For example, personal computers, tablet computers, smart phones, personal digital assistants (PDAs), game consoles, Internet Protocol Televisions (IPTVs), smart wearable devices, etc.

[0033] The computer device can also include network devices and / or user devices. Among them, the network devices include but are not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0034] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0035] Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0036] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0037] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0038] S10, in response to the data synchronization instruction from the source database to the target database, construct and publish a configuration file to obtain a consumer, a distributor, a complement operator, and an outputter.

[0039] In this embodiment, the source database refers to the database that needs to synchronize data to the downstream business system. For example, in the financial field, the source database can be an exchange, etc.

[0040] In this embodiment, the target database can be the destination end of data synchronization, which is used to receive the data sent from the source database. For example, in the financial field, the target database can be a trading system.

[0041] In this embodiment, the constructing and publishing the configuration file includes:

[0042] Configure source information and sink information in the BS (Browser / Server Architecture) architecture;

[0043] Among them, the source information includes the IP (Internet Protocol) port of the source database, the database name, the table name to be synchronized, the username of the database, and the password of the database;

[0044] Among them, the sink information includes the IP port of the target database, the database name, the username of the database, and the password of the database;

[0045] Among them, when the data to be synchronized is not directly obtained from the source database, distributed stream processing platform information is also configured in the BS architecture; the data to be synchronized in the source database is synchronized to the distributed stream processing platform through the distributed stream processing platform information, and the data to be synchronized is obtained from the distributed stream processing platform.

[0046] For example: If the control requirements state that the data to be synchronized must first enter the distributed stream processing platform for centralized extraction, and then the data in the distributed stream processing platform is extracted as the source data for downstream synchronization, then configure the information of the distributed stream processing platform; if there are no special control requirements, there is no need to configure the information of the distributed stream processing platform.

[0047] Through the above embodiments, the distributed stream processing platform can be flexibly configured according to actual synchronization requirements, thereby effectively saving deployment costs.

[0048] Moreover, only simple configuration and publishing are required to execute subsequent non-full log supplementation operations and data synchronization operations, without the need for users to write code, effectively reducing labor costs.

[0049] S11. Obtain the table structure information (base schema) of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain the target table structure of the target database.

[0050] For example: First, obtain the table structure information from the source database, such as table names, all fields of the table, and properties information (attribute information), and then use a set of programs to combine the information. Specifically:

[0051] For the table user_info., there are 2 fields in the table: id int, name varchar.

[0052] You can call a specified command to obtain this basic information, and then splice together the table creation statement:

[0053] CREATE TABLE IF NOT EXISTS user_info(

[0054] Id int,

[0055] Name varchar );

[0057] Furthermore, execute this statement in the target database to achieve table structure initialization and obtain the target table structure of the target database.

[0058] S12. Use the consumer to consume the log data in the source database.

[0059] In this embodiment, using the consumer to consume the log data in the source database includes:

[0060] Use the consumer to obtain data from the source database and convert the obtained data into the target table structure.

[0061] For example, data can be consumed line by line to achieve ordered consumption of the data and avoid anomalies.

[0062] Among them, if the distributed stream processing platform exists, log data can be read line by line from the distributed stream processing platform in consumer mode, and the obtained data can be converted into the target table structure to conform to the input format of the downstream database.

[0063] Through the above embodiments, the log data to be synchronized can be obtained from the source database.

[0064] S13. Use the distributor to detect the log type of the log data, and write the log data into a pre-constructed original array or padding array according to the log type.

[0065] In this embodiment, the log type may include full log type and non-full log type.

[0066] Among them, the full log type refers to a log without missing fields.

[0067] Among them, the non-full log type refers to a log with missing fields.

[0068] In this embodiment, writing the log data into a pre-constructed original array or padding array according to the log type includes:

[0069] When the log type is insert type (insert) or delete type (del), use the distributor to write the log data into the padding array; or

[0070] When the log type is update type (update), use the distributor to write the log data into the original array.

[0071] Among them, the data contained in the padding array has no missing fields and belongs to data that does not need to be padded or has already been padded.

[0072] Among them, the data contained in the original array has missing fields and belongs to data that needs to be padded.

[0073] The padding array and the original array can be custom-configured and stored in memory for easy calling.

[0074] S14. Use the distributor to asynchronously start the padding operator, and use the padding operator to perform padding processing on each piece of data in the original array, and write each piece of data obtained after the padding processing into the padding array.

[0075] In this embodiment, the asynchronous startup of the complement operator by the distributor and the use of the complement operator to perform complement processing on each piece of data in the original array include:

[0076] When it is detected that log data of the update type is written into the original array, the distributor is used to asynchronously start the complement operator; wherein, each piece of log data of the update type corresponds to a complement operator;

[0077] Use the corresponding complement operator to obtain the data ID (Identity document) of each piece of data in the original array;

[0078] Use the corresponding complement operator to query for missing fields in the historical data cached locally according to the data ID of each piece of data;

[0079] When the missing fields are not found in the historical data cached locally, use the corresponding complement operator to remotely query for missing fields in the historical data of the downstream engine according to the data ID of each piece of data;

[0080] When the missing fields are not found in the historical data of the downstream engine, use the corresponding complement operator to query for missing fields in the upstream system according to the data ID of each piece of data;

[0081] Use the corresponding complement operator to fill the queried missing fields into each piece of data in the original array.

[0082] Among them, the complement operator is asynchronous with the entire consumption link, that is, concurrent processing, which effectively reduces the possibility of blocking and thus improves the processing efficiency.

[0083] For example: For the data X in the original array: {id: 1, age: 15}, the data X is missing a field name;

[0084] First, query according to the id of data X in the local cache (Cache_local). Since the speed based on memory is very fast, efficient query can be guaranteed;

[0085] If the missing fields of data X are found in Cache_local, directly read the data in cache_local ({id, 1, age: 14: name: Lily}), and then complete it to: {id: 1, age: 15, name: Lily}.

[0086] If the missing fields of data X are not found in Cache_local, then remote query the data in the downstream engine, select * from user_info where id = 1; retrieve the historical data and then complete it as: {id: 1, age: 15, name: Lily}.

[0087] Among them, the field name: Lily is the missing field retrieved.

[0088] If the downstream engine also fails to find it, continue to query in the upstream system, and wait until the missing fields are retrieved and then complete the filling.

[0089] After the completion of the completion operation, the completed result {id: 1, age: 15, name: Lily} data can be placed in the local cache and the completion array.

[0090] Through the above embodiments, non-full logs can be automatically and asynchronously completed, and concurrent operations can also avoid data blockage, thereby ensuring the accuracy of data synchronization; moreover, the method of gradually querying missing fields further ensures the completion efficiency.

[0091] At the same time, when there is a change in the table structure, such as adding a certain field, the change can be automatically captured, and it can be automatically adapted and synchronously modified without the user having to rebuild the synchronization program.

[0092] S15, start the outputter according to the start strategy, and use the outputter to merge the data in the completion array to obtain the target data, and send the target data to the target database.

[0093] In this embodiment, the starting the outputter according to the start strategy includes:

[0094] When it is detected that the data in the completion array reaches the preset data volume, start the outputter;

[0095] When it is detected that the time duration since the last start of the outputter reaches the configured duration, start the outputter.

[0096] For example: the preset quantity can be 10,000 pieces, and the configured duration can be 1 minute. In this way, as long as the data in the completion array reaches 10,000 pieces, the outputter is started for data merging to avoid data overload; if the data in the completion array does not reach 10,000 pieces when the time duration since the last start of the outputter reaches 1 minute, the outputter is still started at this time to perform data merging and sending in a timely manner, ensuring the data synchronization efficiency and avoiding ineffective waiting.

[0097] In this embodiment, the step of using the outputter to merge the data in the supplemented array to obtain target data includes:

[0098] For multiple pieces of data in the supplemented array that have the same data ID, retain the latest fields of each piece of data in the multiple pieces of data to obtain the target data.

[0099] For example: The data composition of each log data is (id, age, name). The original array and the supplemented array are shown in the following table:

[0100] Original array 1,15 2,18 4,30 4,31 Complemented array 1, 15, Lily 2, 16, Tom 2, 18, Tom 3, del 4, 30, John 4, 31, John

[0101] For the data with id = 1, after merging according to the latest fields, the target data (1, 15, Lily) can be obtained; for the data with id = 2, after merging according to the latest fields, the target data (2, 18, Tom) can be obtained; for the data with id = 3, after merging according to the latest fields, the target data (3, del) can be obtained; for the data with id = 4, after merging according to the latest fields, the target data (4, 31, John) can be obtained.

[0102] Through the above embodiments, it is possible to merge data before sending it downstream, so as to reduce the sending of unnecessary data and improve the writing efficiency.

[0103] In the financial field, through the completion and synchronization of non-full logs, it is possible to read non-full logs during data synchronization and perform automatic supplementation. Users do not need to write a large amount of configuration information, nor write code or execute statements. Moreover, when the upstream table changes, it can be automatically recognized and the changes can be synchronized to the downstream, effectively improving the efficiency of data synchronization between databases in the financial field.

[0104] From the above technical solutions, on the one hand, constructing and publishing the configuration file to obtain the consumer, distributor, supplementation operator, and outputter realizes code-free operation and effectively reduces the labor cost; on the other hand, the distributor asynchronously starts the supplementation operator, and the supplementation operator performs supplementation processing on each piece of data in the original array, which can realize the automatic supplementation of non-full logs during data synchronization, thereby ensuring the accuracy of data synchronization. Moreover, concurrent operations can also achieve the rapid supplementation of log data and avoid data blocking; on the other hand, using the outputter to merge the data in the supplemented array and then send it to the target database can avoid sending unnecessary data to the target database and improve the data writing efficiency.

[0105] Such as Figure 2As shown in the figure, it is a functional block diagram of a preferred embodiment of the data synchronization device based on non-full log completion according to the present invention. The data synchronization device 11 based on non-full log completion includes a publishing unit 110, an initialization unit 111, a consuming unit 112, a writing unit 113, a completion unit 114, and a merging unit 115. The module / unit referred to in the present invention means a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0106] The publishing unit 110 is configured to construct and publish a configuration file in response to a data synchronization instruction from a source database to a target database, and obtain a consumer, a distributor, a completion operator, and an outputter.

[0107] In this embodiment, the source database refers to a database that needs to synchronize data to a downstream business system. For example: in the financial field, the source database can be an exchange, etc.

[0108] In this embodiment, the target database can be the destination end of data synchronization, and is used to receive data sent from the source database. For example: in the financial field, the target database can be a trading system.

[0109] In this embodiment, the publishing unit 110 constructing and publishing a configuration file includes:

[0110] Configure source information and sink information in the BS (Browser / Server Architecture) architecture;

[0111] Among them, the source information includes the IP (Internet Protocol) port of the source database, the database name, the table name to be synchronized, the username of the database, and the password of the database;

[0112] Among them, the sink information includes the IP port of the target database, the database name, the username of the database, and the password of the database;

[0113] Among them, when the data to be synchronized is not directly obtained from the source database, distributed stream processing platform information is also configured in the BS architecture; the data to be synchronized in the source database is synchronized to the distributed stream processing platform through the distributed stream processing platform information, and the data to be synchronized is obtained from the distributed stream processing platform.

[0114] For example: If the control requirements state that the data to be synchronized must first enter the distributed stream processing platform for centralized extraction, and then the data in the distributed stream processing platform is extracted as the source data for downstream synchronization, then configure the information of the distributed stream processing platform; if there are no special control requirements, there is no need to configure the information of the distributed stream processing platform.

[0115] Through the above embodiments, the distributed stream processing platform can be flexibly configured according to actual synchronization requirements, thereby effectively saving deployment costs.

[0116] Moreover, only simple configuration and publication are required to execute subsequent non-full log supplementation operations and data synchronization operations, without the need for users to write code, effectively reducing labor costs.

[0117] The initialization unit 111 is used to obtain the table structure information (base schema) of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain the target table structure of the target database.

[0118] For example: First, obtain the table structure information from the source database, such as table names, all fields of the table, and properties information (attribute information), and then combine the information with a set of programs. Specifically:

[0119] For the table user_info., there are 2 fields in the table: id int, name varchar.

[0120] You can call a specified command to obtain this basic information, and then splice out the table creation statement:

[0121] CREATE TABLE IF NOT EXISTS user_info(

[0122] Id int,

[0123] Name varchar );

[0125] Furthermore, execute this statement in the target database to achieve table structure initialization and obtain the target table structure of the target database.

[0126] The consumption unit 112 is used to consume the log data in the source database by using the consumer.

[0127] In this embodiment, the consumption unit 112 consuming the log data in the source database by using the consumer includes:

[0128] Using the consumer to obtain data from the source database and convert the obtained data into the target table structure.

[0129] For example, data can be consumed line by line to achieve ordered consumption of data and avoid anomalies.

[0130] Among them, if there is the distributed stream processing platform, the log data can be read line by line from the distributed stream processing platform in consumer mode, and the obtained data can be converted into the target table structure to conform to the input format of the downstream database.

[0131] Through the above embodiments, the log data to be synchronized can be obtained from the source database.

[0132] The writing unit 113 is configured to use the distributor to detect the log type of the log data, and write the log data into a pre-constructed original array or padded array according to the log type.

[0133] In this embodiment, the log type may include a full log type and a non-full log type.

[0134] Among them, the full log type refers to a log without missing fields.

[0135] Among them, the non-full log type refers to a log with missing fields.

[0136] In this embodiment, the writing unit 113 writing the log data into a pre-constructed original array or padded array according to the log type includes:

[0137] When the log type is the insert type (insert) or the delete type (del), use the distributor to write the log data into the padded array; or

[0138] When the log type is the update type (update), use the distributor to write the log data into the original array.

[0139] Among them, the data included in the padded array has no missing fields and belongs to data that does not need to be padded or has already been padded.

[0140] Among them, the data included in the original array has missing fields and belongs to data that needs to be padded.

[0141] The padded array and the original array can be custom-configured and stored in memory for convenient invocation.

[0142] The complementing unit 114 is configured to use the distributor to asynchronously start the padding operator, and use the padding operator to perform a complementing process on each piece of data in the original array, and write each piece of data obtained after the complementing process into the padded array.

[0143] In this embodiment, the completion unit 114 asynchronously starts the completion operator by using the distributor, and the process of using the completion operator to complete each piece of data in the original array includes:

[0144] When it is detected that log data of the update type is written into the original array, the distributor is used to asynchronously start the completion operator; wherein, each piece of log data of the update type corresponds to a completion operator;

[0145] Use the corresponding completion operator to obtain the data ID (Identity document) of each piece of data in the original array;

[0146] Use the corresponding completion operator to query for missing fields in the historical data cached locally according to the data ID of each piece of data;

[0147] When the missing fields are not found in the historical data cached locally, use the corresponding completion operator to remotely query for missing fields in the historical data of the downstream engine according to the data ID of each piece of data;

[0148] When the missing fields are not found in the historical data of the downstream engine, use the corresponding completion operator to query for missing fields in the upstream system according to the data ID of each piece of data;

[0149] Use the corresponding completion operator to fill the queried missing fields into each piece of data in the original array.

[0150] Among them, the completion operator is asynchronous with the entire consumption link, that is, concurrent processing, which effectively reduces the possibility of blocking, thereby improving the processing efficiency.

[0151] For example: For the data X in the original array: {id: 1, age: 15}, the data X is missing a field name;

[0152] First, query according to the id of the data X in the local cache (Cache_local). Since the speed based on memory is very fast, efficient query can be guaranteed;

[0153] If the missing fields of the data X are found in the Cache_local, directly read the data in the cache_local ({id, 1, age: 14: name: Lily}), and then complete it to: {id: 1, age: 15, name: Lily}.

[0154] If the missing fields of data X are not found in Cache_local, the data will be remotely queried in the downstream engine, select * from user_info where id = 1; the historical data will be retrieved and then completed to: {id: 1, age: 15, name: Lily}.

[0155] Among them, the field name: Lily is the missing field retrieved.

[0156] If the downstream engine also fails to find it, continue to query in the upstream system, and wait until the missing fields are retrieved and then complete the filling.

[0157] After the completion of the filling operation, the filled result {id: 1, age: 15, name: Lily} data can be placed in the local cache and the filling array.

[0158] Through the above embodiments, non-full logs can be automatically and asynchronously filled, and concurrent operations can also avoid data blockage, thus ensuring the accuracy of data synchronization; moreover, the method of gradually querying missing fields further guarantees the filling efficiency.

[0159] Meanwhile, when there is a change in the table structure, such as adding a certain field, the change can be automatically captured, and it can be automatically adapted and synchronously modified without the user having to reconstruct the synchronization program.

[0160] The merging unit 115 is configured to start the outputter according to the start strategy, and use the outputter to merge the data in the filling array to obtain target data, and send the target data to the target database.

[0161] In this embodiment, the merging unit 115 starting the outputter according to the start strategy includes:

[0162] When it is detected that the data in the filling array reaches the preset data volume, start the outputter;

[0163] When it is detected that the time duration since the last start of the outputter reaches the configured time duration, start the outputter.

[0164] For example: the preset quantity can be 10,000 pieces, and the configured time duration can be 1 minute. In this way, as long as the data in the filling array reaches 10,000 pieces, the outputter will be started for data merging to avoid data overload; if the data in the filling array does not reach 10,000 pieces when the time duration since the last start of the outputter reaches 1 minute, the outputter will still be started at this time to perform data merging and sending in a timely manner, ensuring data synchronization efficiency and avoiding ineffective waiting.

[0165] In this embodiment, the merging unit 115 uses the outputter to merge the data in the padded array to obtain target data, which includes:

[0166] For multiple pieces of data in the padded array that have the same data ID, retain the latest fields of each piece of data in the multiple pieces of data to obtain the target data.

[0167] For example: The data composition of each log data is (id, age, name). The original array and the padded array are shown in the following table:

[0168] Original array 1,15 2,18 4,30 4,31 Complemented array 1, 15, Lily 2, 16, Tom 2, 18, Tom 3, del 4, 30, John 4, 31, John

[0169] For the data with id = 1, after merging according to the latest fields, the target data (1, 15, Lily) can be obtained; for the data with id = 2, after merging according to the latest fields, the target data (2, 18, Tom) can be obtained; for the data with id = 3, after merging according to the latest fields, the target data (3, del) can be obtained; for the data with id = 4, after merging according to the latest fields, the target data (4, 31, John) can be obtained.

[0170] Through the above embodiments, it is possible to merge data before sending it downstream, so as to reduce the sending of unnecessary data and improve the writing efficiency.

[0171] In the financial field, through the completion and synchronization of non-full logs, it is possible to read non-full logs during data synchronization and perform automatic padding. Users do not need to write a large amount of configuration information, do not need to write code or execute statements, and can automatically identify when the upstream table changes and synchronize the changes to the downstream, effectively improving the efficiency of data synchronization between databases in the financial field.

[0172] From the above technical solutions, on the one hand, constructing and publishing the configuration file to obtain the consumer, distributor, padding operator, and outputter realizes code-free, effectively reducing the labor cost; on the other hand, the distributor asynchronously starts the padding operator, and the padding operator performs padding processing on each piece of data in the original array, which can realize the automatic padding of non-full logs during data synchronization, thereby ensuring the accuracy of data synchronization. Moreover, concurrent operations can also achieve fast padding of log data and avoid data blocking; on the other hand, using the outputter to merge the data in the padded array and then send it to the target database can avoid sending unnecessary data to the target database and improve the data writing efficiency.

[0173] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment of the data synchronization method based on non-full log padding of the present invention.

[0174] The computer device 1 may include a memory 12, a processor 13, and a bus. It may also include a computer program stored in the memory 12 and executable on the processor 13, such as a data synchronization program based on non-full log completion.

[0175] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus-type structure or a star structure. The computer device 1 may also include more or fewer other hardware or software components than those shown in the figure, or different component arrangements. For example, the computer device 1 may also include input / output devices, network access devices, etc.

[0176] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are hereby incorporated by reference.

[0177] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as the mobile hard disk of the computer device 1. In other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both an internal storage unit and an external storage device of the computer device 1. The memory 12 can be used not only to store application software installed on the computer device 1 and various types of data, such as the code of a data synchronization program based on non-full log completion, but also to temporarily store data that has been output or will be output.

[0178] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and circuits. By running or executing programs or modules stored in the memory 12 (such as executing a data synchronization program based on non-full log completion), and by calling data stored in the memory 12, it performs various functions of the computer device 1 and processes data.

[0179] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of the data synchronization method based on non-full log completion, such as Figure 1 the steps shown.

[0180] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a publishing unit 110, an initialization unit 111, a consumption unit 112, a writing unit 113, a completion unit 114, and a merging unit 115.

[0181] The above-mentioned integrated units implemented in the form of software function modules may be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the data synchronization method based on non-full log completion described in various embodiments of the present invention.

[0182] If the modules / units integrated in the computer device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented.

[0183] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory, etc.

[0184] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.

[0185] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0186] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is set to realize the connection and communication between the memory 12 and at least one processor 13, etc.

[0187] Although not shown, the computer device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source may be logically connected to the at least one processor 13 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The computer device 1 may further include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0188] Further, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the computer device 1 and other computer devices.

[0189] Optionally, the computer device 1 may further include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.

[0190] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0191] Those skilled in the art can understand that Figure 3 the shown structure does not limit the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0192] Combined with Figure 1 , the memory 12 in the computer device 1 stores a plurality of instructions to implement a data synchronization method based on non-full log completion. The processor 13 can execute the plurality of instructions to implement:

[0193] In response to a data synchronization instruction from a source database to a target database, construct and publish a configuration file to obtain a consumer, a distributor, a completion operator, and an outputter;

[0194] Obtain the table structure information of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain the target table structure of the target database;

[0195] Use the consumer to consume the log data in the source database;

[0196] Use the distributor to detect the log type of the log data, and write the log data into a pre-constructed original array or padding array according to the log type;

[0197] Use the distributor to asynchronously start the padding operator, and use the padding operator to perform padding processing on each piece of data in the original array, and write each piece of data obtained after the padding processing into the padding array;

[0198] Start the outputter according to the start strategy, and use the outputter to merge the data in the padding array to obtain the target data, and send the target data to the target database.

[0199] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0200] It should be noted that all the data involved in this case are legally obtained. The non-company software tools or components appearing in the embodiments of this application are only for example introduction and do not represent actual use.

[0201] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0202] The present invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0203] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0205] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0206] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0207] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. The terms first, second, etc. are used to denote names and do not denote any specific order.

[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A data synchronization method based on non-full log completion, characterized in that: The data synchronization method based on non-full log completion includes: In response to the data synchronization instruction from the source database to the target database, the configuration file is constructed and published to obtain the consumer, the distributor, the completion operator and the outputter; Acquire table structure information of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain a target table structure of the target database; Using the consumer to consume the log data in the source database; Detecting the log type of the log data using the distributor, and writing the log data into a pre-built original array or a padded array according to the log type; Using the distributor to asynchronously start the padding operator, and using the padding operator to perform padding processing on each piece of data in the original array, and writing each piece of data obtained after the padding processing into the padding array; The output device is started according to the start-up strategy, and the data in the padded array is merged by the output device to obtain target data, and the target data is sent to the target database.

2. The data synchronization method based on non-full log completion according to claim 1, characterized in that: The build and publish configuration files include: Configure source information and sink information in the BS architecture; The source information includes the IP port of the source database, the database name, the table name to be synchronized, the database user name, and the database password; The sink information includes the IP port of the target database, the database name, the database user name, and the database password; Among them, when the data to be synchronized is not directly obtained from the source database, distributed stream processing platform information is also configured in the BS architecture; the data to be synchronized in the source database is synchronized to the distributed stream processing platform through the distributed stream processing platform information, and the data to be synchronized is obtained from the distributed stream processing platform.

3. The data synchronization method based on non-full log completion according to claim 1, characterized in that: The using the consumer to consume the log data in the source database includes: The consumer is used to obtain data from the source database, and the obtained data is converted into the target table structure.

4. The data synchronization method based on non-full log completion according to claim 1, characterized in that: Writing the log data into a pre-built original array or a padded array according to the log type includes: When the log type is an insertion type or a deletion type, using the distributor to write the log data into the padding array; or When the log type is an update type, the log data is written into the original array using the distributor.

5. The data synchronization method based on non-full log completion according to claim 4, characterized in that: The step of asynchronously starting the padding operator by using the distributor, and performing padding processing on each piece of data in the original array by using the padding operator includes: When it is detected that log data of the update type is written into the original array, the distributor is used to asynchronously start the completion operator; wherein each log data of the update type corresponds to a completion operator; Obtain the data ID of each piece of data in the original array using the corresponding padding operator; Use the corresponding completion operator to query the missing fields in the locally cached historical data according to the data ID of each piece of data; When the missing field is not found in the historical data of the local cache, the missing field is remotely queried in the historical data of the downstream engine according to the data ID of each piece of data using the corresponding completion operator; When the missing fields are not found in the historical data of the downstream engine, the corresponding completion operator is used to query the missing fields in the upstream system according to the data ID of each data record; The missing fields found in the search are filled into each piece of data in the original array using the corresponding filling operator.

6. The data synchronization method based on non-full log completion according to claim 1, characterized in that: The starting of the output device according to the starting strategy includes: When it is detected that the data in the padding array reaches a preset data amount, starting the output device; When it is detected that the time from the last start of the output device reaches the configured time, the output device is started.

7. The data synchronization method based on non-full log completion according to claim 1, characterized in that: The step of merging the data in the padding array using the output device to obtain the target data comprises: For multiple pieces of data with the same data ID in the padding array, the latest field of each piece of data in the multiple pieces of data is retained to obtain the target data.

8. A data synchronization device based on non-full log completion, characterized in that: The data synchronization device based on incomplete log completion includes: A publishing unit, for responding to a data synchronization instruction from a source database to a target database, constructing and publishing a configuration file, and obtaining a consumer, a distributor, a completion operator, and an outputter; An initialization unit, used to obtain table structure information of the source database, and perform table structure initialization processing in the target database according to the table structure information to obtain a target table structure of the target database; A consumption unit, configured to consume the log data in the source database using the consumer; A writing unit, configured to detect a log type of the log data using the distributor, and write the log data into a pre-built original array or a padded array according to the log type; A completion unit, used to asynchronously start the completion operator using the distributor, and to complete each piece of data in the original array using the completion operator, and to write each piece of data obtained after the completion process into the completion array; A merging unit is used to start the output device according to the startup strategy, and use the output device to merge the data in the padded array to obtain target data, and send the target data to the target database.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the data synchronization method based on non-full log completion as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the data synchronization method based on non-full log completion as described in any one of claims 1 to 7.