Heterogeneous data synchronization method and device and storage medium
By reading preset configuration files and building a heterogeneous data synchronization environment, the problem of difficulty in writing database statements under multiple data sources is solved, efficient and flexible heterogeneous data synchronization is achieved, and the operation process is simplified and ease of use and scalability is improved.
Patent Information
- Application Number
- CN202311830717.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
When multiple data sources exist, it is difficult to write corresponding database statements, which requires a certain technical foundation for operators, and there are problems of operational difficulties.
By reading preset configuration files, dynamically configure the heterogeneous data sets, field mapping relationships, deduplication rules and data sources that need to be synchronized, to build a heterogeneous data synchronization environment, and summarize the heterogeneous data sets into the target data set through field deduplication and field mapping processing.
No need to manually write complex database statements, improve the efficiency and flexibility of heterogeneous data synchronization, simplify operational processes, improve ease of use, and enhance the scalability of heterogeneous data synchronization.
Smart Images

Figure CN120234366A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a heterogeneous data synchronization method, apparatus, and storage medium. Background Art
[0002] With the development of mobile Internet and Internet of Things technologies, data sources have become increasingly diverse, posing great challenges to data normalization and summarization. During the data summarization process, it is often encountered that there are multiple branch units under a single entity, and each unit has different groups; or there are different cities, counties, and townships in a province, and each township has many entities. Collecting the statistical data of these entities or extracting data from the databases maintained by each entity often requires data summarization according to a unified standard. In the real world, this is a time-consuming and labor-intensive process. However, synchronizing and summarizing heterogeneous data generated from multiple channels according to a unified standard is a common task scenario.
[0003] Traditional heterogeneous data synchronization methods include: performing insert, update, delete, and query operations on heterogeneous data in different data sources through database statements; extracting heterogeneous data from each different data source, and sequentially cleaning and transforming each heterogeneous data before storing it in the target database.
[0004] However, in the case of multiple data sources, it is difficult to write the corresponding database statements, which requires a certain technical foundation for operators and poses difficulties in operation. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a heterogeneous data synchronization method, apparatus, and storage medium to eliminate or improve one or more defects existing in the prior art, and can solve the problem that it is difficult to write the corresponding database statements in the case of multiple data sources, which requires a certain technical foundation for operators and poses difficulties in operation.
[0006] One aspect of the present invention provides a heterogeneous data synchronization method, including:
[0007] Reading a preset configuration file; the preset configuration file includes several heterogeneous data sets that need data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets;
[0008] Based on the preset configuration file, constructing a heterogeneous data synchronization environment;
[0009] Performing field deduplication processing on the heterogeneous data sets based on the field deduplication rules;
[0010] Perform field mapping processing on the heterogeneous data set through the mapping relationship and the preset field processing function corresponding to the heterogeneous data set, and summarize and store the heterogeneous data set in the target data set.
[0011] Optionally, perform field deduplication processing on the heterogeneous data set based on the field deduplication rule, including: performing at least one round of deduplication processing on the heterogeneous data set based on the field deduplication rule, and performing data persistence processing on the heterogeneous data set after each round of deduplication processing.
[0012] Optionally, performing data persistence processing on the heterogeneous data set after each round of deduplication processing includes: storing the heterogeneous data set after each round of deduplication processing in the target database; or storing the heterogeneous data set after each round of deduplication processing in the form of a file; or storing the heterogeneous data set after each round of deduplication processing in the form of a log in the log database; or storing the heterogeneous data set after each round of deduplication processing in the target cache.
[0013] Optionally, each round of deduplication processing performed on the heterogeneous data set includes internal deduplication processing and overall deduplication processing.
[0014] Optionally, the field deduplication rule includes a single-field deduplication rule and / or a multi-field deduplication rule.
[0015] Optionally, the heterogeneous data set includes a structured data set and an unstructured data set; before performing field deduplication processing on the heterogeneous data set based on the field deduplication rule, it further includes: preprocessing the heterogeneous data set, and in the case where the heterogeneous data set is a structured data set, the preprocessing includes data deduplication processing and special character removal processing.
[0016] Optionally, in the case where the heterogeneous data set is an unstructured data set, the preprocessing includes information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set.
[0017] Optionally, information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set includes: identifying and extracting entities in the unstructured data set through a preset entity recognition model; or matching the unstructured data set through a preset regular expression for information extraction; or analyzing and processing the unstructured data set through natural language processing technology to achieve information extraction.
[0018] Another aspect of the present invention provides a device for heterogeneous data synchronization, the device includes: a processor and a memory, characterized in that computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the device implements the steps of the above heterogeneous data synchronization method.
[0019] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the above heterogeneous data synchronization method are implemented.
[0020] The beneficial effects of the present invention are at least as follows:
[0021] The heterogeneous data synchronization method, device and storage medium of the present invention read a preset configuration file; the preset configuration file includes several heterogeneous data sets that need data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets; based on the preset configuration file, a heterogeneous data synchronization environment is constructed; based on the field deduplication rules, field deduplication processing is performed on the heterogeneous data sets; through the mapping relationship and the preset field processing function corresponding to the heterogeneous data set, field mapping processing is performed on the heterogeneous data set, and the heterogeneous data sets are summarized and stored in the target data set. It can solve the problem that it is relatively difficult to write corresponding database statements in the case of multiple data sources, which requires a certain technical foundation for operators and there are operation difficulties. By reading the preset configuration file, information such as heterogeneous data sets to be synchronized, field mapping relationships, deduplication rules, and data sources can be dynamically configured. In this way, it can adapt to different data sources and data requirements without manually writing complex database statements, improving the efficiency and flexibility of heterogeneous data synchronization; at the same time, storing the configuration information of data synchronization in the preset configuration file separates the configuration information from the code logic. In this way, when the requirements change or new data sources are added, only the configuration file needs to be updated without modifying the code. Through the preset configuration file, heterogeneous data sets to be synchronized can be added or deleted, and the field mapping relationship and deduplication rules can be flexibly adjusted, which can easily handle the access of new data sources or the change of old data sources, improving the scalability of heterogeneous data synchronization.
[0022] Furthermore, by using the preset field processing function and the field mapping relationship, the fields of the heterogeneous data set can be mapped to the standard fields of the target data set. In this way, operators do not need to deeply understand database statements and data conversion logic, and only need to perform simple configuration according to the preset configuration file to achieve data synchronization operations, improving the usability of heterogeneous data synchronization.
[0023] The additional advantages, objectives, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the description and the drawings.
[0024] Those skilled in the art will understand that the objectives and advantages achievable by the present invention are not limited to the above particulars, and the above and other objectives achievable by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0026] Figure 1 It is a flowchart of a heterogeneous data synchronization method according to an embodiment of the present invention;
[0027] Figure 2 It is a block diagram of a heterogeneous data synchronization device according to an embodiment of the present invention;
[0028] Figure 3 It is a block diagram of a heterogeneous data synchronization device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0030] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0031] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0032] Here, it should also be noted that if not otherwise specified, the term "connection" in this article can not only refer to a direct connection, but also represent an indirect connection with an intermediate.
[0033] In the following, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0034] The heterogeneous data synchronization method provided by this application will be introduced in detail below.
[0035] As Figure 1As shown in the figure, an embodiment of the present application provides a heterogeneous data synchronization method. This embodiment is described by taking the method being used in an electronic device as an example, where the electronic device includes but is not limited to mobile phones, computers, servers, etc. Specifically, the heterogeneous data synchronization method at least includes the following steps S101 to S104:
[0036] Step S101: Read a preset configuration file.
[0037] In this embodiment, the preset configuration file includes several heterogeneous data sets that require data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets.
[0038] Among them, several heterogeneous data sets include data sets composed of data of various different types, formats, data sources, or semantics. For example: database table data from a database; various format files such as text files, log files, configuration files, XML files, JSON files, etc. from a file system; environmental data or industrial data collected from sensor devices; network data, including data crawled from websites, data obtained from API interfaces, social media data, cloud storage data, etc.; data collected in a manual form. This embodiment does not limit the collection method of several heterogeneous data sets.
[0039] In this embodiment, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields is a pre-constructed mapping relationship.
[0040] For the fields in several heterogeneous data sets, standardized and unified processing is required, including: obtaining the data structures of each heterogeneous data set; determining the common fields and non-common fields based on the data structures; performing induction processing on the non-common fields, and generating a data unification standard in combination with the common fields.
[0041] Among them, the data unification standard includes data format, data structure, data naming convention, data encoding and character set, data value range convention, or data quality standard, etc.
[0042] After obtaining the data unified annotation, according to the business rules and the data unification standard after synchronization of the heterogeneous data sets, determine the synchronization standard fields corresponding to the target data set after synchronization, and use a standard naming format to standardize the field names for subsequent data aggregation of the heterogeneous data sets. Then, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields is a pre-constructed mapping relationship. Sequentially, according to the characteristics of the data structures of each heterogeneous data set, and compare with the synchronization standard fields to formulate the field mapping relationship.
[0043] Specifically, establishing the mapping relationship between the fields in the heterogeneous dataset and the synchronization standard fields includes: determining the field information of different heterogeneous datasets; the field information includes data structure, field meaning, or data type; based on the field information of different heterogeneous datasets, determining the common fields and non-common fields in each heterogeneous dataset; according to business requirements and the goal of data synchronization, determining the synchronization standard fields, where the synchronization standard fields can cover the necessary information in each heterogeneous dataset and have a unified naming specification; comparing the fields in each heterogeneous dataset with the synchronization standard fields to determine the similarities and differences between the fields in each heterogeneous dataset and the synchronization standard fields, obtaining the comparison result; establishing the field mapping relationship in the configuration file according to the comparison result, and corresponding the heterogeneous dataset fields with the synchronization standard fields one by one.
[0044] In addition, it should be noted that establishing the field mapping relationship is an iterative process. In actual implementation, some special situations and complexities may be encountered, and adjustments and flexible handling need to be made according to the actual situation. Based on this, after establishing the field mapping relationship in the configuration file according to the comparison result, it also includes: verifying the field mapping relationship to obtain the verification result; adjusting the field mapping relationship based on the verification result.
[0045] In this embodiment, in order to ensure the uniqueness of the records or entries in the target dataset after synchronization, avoid duplicate data in the dataset, and at the same time ensure the accuracy and integrity of the data in the target dataset, the field deduplication rules for the heterogeneous dataset are set in the preset configuration file.
[0046] Optionally, the field deduplication rules include single-field deduplication rules and / or multi-field deduplication rules.
[0047] Step S102, based on the preset configuration file, construct a heterogeneous data synchronization environment.
[0048] In this embodiment, the data source corresponding to the heterogeneous dataset includes a database, and the preset configuration file also includes the database information corresponding to the heterogeneous dataset, including the database address, username, and user password; at the same time, the preset configuration file also includes the target database information corresponding to the target database, including the target database address, username, and user password, and the target database is used to store the target dataset after heterogeneous data synchronization.
[0049] Based on this, constructing a heterogeneous data synchronization environment based on the preset configuration file includes: establishing a connection with the database corresponding to the heterogeneous dataset according to the database information in the preset configuration file; establishing a connection with the target database according to the target database information; in the case where the target dataset does not exist in the target database, creating the table structure corresponding to the target dataset according to the synchronization standard fields in the preset configuration file.
[0050] Step S103: perform field deduplication on the heterogeneous data set based on the field deduplication rule.
[0051] In this embodiment, in order to ensure the quality of the synchronized data and improve the reliability of the synchronized data, at least two rounds of processing are required during the field deduplication process of the heterogeneous data set, and data persistence processing is performed after each round of processing.
[0052] Optionally, performing field deduplication on the heterogeneous data set based on the field deduplication rule includes: performing at least one round of deduplication on the heterogeneous data set based on the field deduplication rule, and performing data persistence processing on the heterogeneous data set after each round of deduplication.
[0053] Among them, performing data persistence processing on the heterogeneous data set after each round of deduplication includes: storing the heterogeneous data set after each round of deduplication into the target database; or storing the heterogeneous data set after each round of deduplication in the form of a file; or storing the heterogeneous data set after each round of deduplication in the form of a log into the log database; or storing the heterogeneous data set after each round of deduplication into the target cache.
[0054] In addition, in order to improve the speed of field deduplication processing and the efficiency of heterogeneous data synchronization, in this embodiment, field deduplication processing is performed by combining internal deduplication processing and overall deduplication processing. In each round of field deduplication processing, first perform field deduplication processing within each heterogeneous data set, and then perform overall field deduplication processing. Specifically, each round of deduplication processing performed on the heterogeneous data set includes internal deduplication processing and overall deduplication processing.
[0055] In addition, in this embodiment, the heterogeneous data includes structured data sets and unstructured data sets, including but not limited to files corresponding to formats such as txt format, csv format, or xlsx format, etc. This embodiment does not limit the file format corresponding to the heterogeneous data set.
[0056] Before performing field deduplication on the heterogeneous data set based on the field deduplication rule, it also includes: preprocessing the heterogeneous data set. In the case where the heterogeneous data set is a structured data set, the preprocessing includes data deduplication processing and special character removal processing. Correspondingly, in the case where the heterogeneous data set is an unstructured data set, the preprocessing includes information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set.
[0057] Among them, information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set includes: identifying and extracting entities in the unstructured data set through a preset entity recognition model; or, performing information extraction by matching the unstructured data set with a preset regular expression; or, analyzing and processing the unstructured data set through natural language processing technology to achieve information extraction.
[0058] Among them, the preset entity recognition model refers to a machine learning model that has been pre-trained to extract specific entities from text, such as the Natural Language Toolkit (NLTK) or the Stanford Named Entity Recognizer (Stanford NER), etc. The entities recognized by the preset entity recognition model can include various types of information such as personal names, place names, organizations, date and time, currency, etc.
[0059] Natural Language Processing (NLP) technology refers to a series of technologies and methods for a computer to understand, process, and generate human language. Such as Part-of-Speech Tagging, Named Entity Recognition, or Word Segmentation and other technologies.
[0060] Step S104, through the mapping relationship and the preset field processing function corresponding to the heterogeneous data set, perform field mapping processing on the heterogeneous data set, and summarize and store the heterogeneous data set into the target data set.
[0061] In this embodiment, the heterogeneous data set includes nested json fields. Based on this, the preset field processing function is a function required for pre-setting the processing of json fields, including a combination of one or more functions among a JSON parsing function, a JSON serialization function, a JSON field extraction function, a JSON field traversal function, a JSON field filtering function, a JSON field merging function, a JSON field sorting function, and a JSON field formatting function.
[0062] The heterogeneous data synchronization method provided in this embodiment reads a preset configuration file. The preset configuration file includes several heterogeneous data sets that require data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets. Based on the preset configuration file, a heterogeneous data synchronization environment is constructed. The fields of the heterogeneous data sets are deduplicated based on the field deduplication rules. Through the mapping relationship and the preset field processing functions corresponding to the heterogeneous data sets, the fields of the heterogeneous data sets are mapped, and the heterogeneous data sets are summarized and stored in the target data set. This can solve the problem that it is difficult to write corresponding database statements in the case of multiple data sources. For operators, certain technical knowledge is required, resulting in operational difficulties. By reading the preset configuration file, information such as the heterogeneous data sets to be synchronized, field mapping relationships, deduplication rules, and data sources can be dynamically configured. In this way, it can adapt to different data sources and data requirements without manually writing complex database statements, improving the efficiency and flexibility of heterogeneous data synchronization. At the same time, storing the configuration information of data synchronization in the preset configuration file separates the configuration information from the code logic. In this way, when the requirements change or new data sources are added, only the configuration file needs to be updated without modifying the code. Through the preset configuration file, heterogeneous data sets to be synchronized can be added or deleted, and the field mapping relationships and deduplication rules can be flexibly adjusted, enabling easy handling of new data source access or changes to old data sources, and improving the scalability of heterogeneous data synchronization.
[0063] Further, by using the preset field processing functions and field mapping relationships, the fields of the heterogeneous data sets can be mapped to the standard fields of the target data set. In this way, operators do not need to deeply understand database statements and data conversion logic. They only need to perform simple configuration according to the preset configuration file to achieve data synchronization operations, improving the usability of heterogeneous data synchronization.
[0064] Figure 2 It is a block diagram of a heterogeneous data synchronization device provided by an embodiment of the present application. The device at least includes the following modules: a file reading module 210, an environment building module 220, a field processing module 230, and a data synchronization module 240.
[0065] The file reading module 210 is used to read the preset configuration file. The preset configuration file includes several heterogeneous data sets that require data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets.
[0066] The environment building module 220 is used to construct a heterogeneous data synchronization environment based on the preset configuration file.
[0067] A field processing module 230 is configured to perform field deduplication processing on the heterogeneous data set based on the field deduplication rule.
[0068] A data synchronization module 240 is configured to perform field mapping processing on the heterogeneous data set through a mapping relationship and a preset field processing function corresponding to the heterogeneous data set, and summarize and store the heterogeneous data set into a target data set.
[0069] For relevant details, refer to the above method embodiments.
[0070] It should be noted that: when the heterogeneous data synchronization device provided in the above embodiments performs heterogeneous data synchronization, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the heterogeneous data synchronization device is divided into different functional modules to complete all or part of the functions described above. In addition, the heterogeneous data synchronization device provided in the above embodiments and the method embodiments of heterogeneous data synchronization belong to the same concept. For the specific implementation process, refer to the method embodiments, which will not be elaborated here.
[0071] This embodiment provides a heterogeneous data synchronization device, as Figure 3 shown, the device at least includes a processor 310 and a memory 320.
[0072] The processor 310 may include one or more processing cores, such as: a 4-core processor, an 8-core processor, etc. The processor 310 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 310 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 310 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 310 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0073] The memory 320 may include one or more computer-readable storage media, which may be non-transitory. The memory 320 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 320 is used to store at least one instruction for being executed by the processor 310 to implement the heterogeneous data synchronization method provided in the method embodiments of the present application.
[0074] In some embodiments, the device may optionally further include: a peripheral device interface and at least one peripheral device. The processor 310, the memory 320, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include, but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0075] Of course, the heterogeneous data synchronization device may further include fewer or more components, and this embodiment does not limit this.
[0076] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the heterogeneous data synchronization method in the above method embodiments.
[0077] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it may be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment may be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0078] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0079] In the present invention, features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0080] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heterogeneous data synchronization method, characterized in that, The method includes: Reading a preset configuration file; the preset configuration file includes several heterogeneous data sets that require data synchronization, the mapping relationship between the fields in the heterogeneous data sets and the synchronization standard fields, the field deduplication rules corresponding to the heterogeneous data sets, and the data sources corresponding to the heterogeneous data sets; Based on the preset configuration file, constructing a heterogeneous data synchronization environment; Performing field deduplication processing on the heterogeneous data sets based on the field deduplication rules; Through the mapping relationship and the preset field processing function corresponding to the heterogeneous data set, performing field mapping processing on the heterogeneous data set, and summarizing and storing the heterogeneous data set into the target data set.
2. The heterogeneous data synchronization method according to claim 1, wherein The performing field deduplication processing on the heterogeneous data sets based on the field deduplication rules includes: Performing at least one round of deduplication processing on the heterogeneous data sets based on the field deduplication rules, and performing data persistence processing on the heterogeneous data sets after each round of deduplication processing.
3. The heterogeneous data synchronization method according to claim 2, wherein The performing data persistence processing on the heterogeneous data sets after each round of deduplication processing includes: Storing the heterogeneous data sets after each round of deduplication processing into the target database; Or, Storing the heterogeneous data sets after each round of deduplication processing in the form of files; Or, Storing the heterogeneous data sets after each round of deduplication processing in the form of logs into the log database; Or, Storing the heterogeneous data sets after each round of deduplication processing into the target cache.
4. The heterogeneous data synchronization method according to claim 2, wherein Each round of deduplication processing performed on the heterogeneous data sets includes internal deduplication processing and overall deduplication processing.
5. The heterogeneous data synchronization method according to claim 1, wherein The field deduplication rules include single-field deduplication rules and / or multi-field deduplication rules.
6. The heterogeneous data synchronization method according to claim 1, wherein The heterogeneous data sets include structured data sets and unstructured data sets; Before performing field deduplication processing on the heterogeneous data sets based on the field deduplication rules, it further includes: Performing preprocessing on the heterogeneous data sets. When the heterogeneous data set is the structured data set, the preprocessing includes data deduplication processing and special character removal processing.
7. The heterogeneous data synchronization method according to claim 6, wherein When the heterogeneous data set is the unstructured data set, the preprocessing includes information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set.
8. The heterogeneous data synchronization method according to claim 7, wherein The information extraction from the unstructured data set based on the business scenario corresponding to the unstructured data set includes: Identifying and extracting entities in the unstructured data set through a preset entity recognition model; Or, Performing information extraction by matching the unstructured data set with a preset regular expression; Or, Performing analysis and processing on the unstructured data set through natural language processing technology to achieve information extraction.
9. An heterogeneous data synchronization device, comprising a processor and a memory, characterized in that, Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.