A method for implementing a data rapid upload engine based on DataX

By designing a data fast upload engine in DataX, using intelligent matching algorithms and visual interfaces to automatically complete the field mapping between data templates and database tables, and dynamically generates json configuration files, the problem of cumbersome configuration of DataX is solved and efficient and fast data upload and import are achieved.

CN115729938BActive Publication Date: 2025-06-17ZHONGBO INFORMATION TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211511821.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-06-17
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

When using DataX for data synchronization in the prior art, the generation of json configuration files is complicated and inefficient, especially in application scenarios of multiple systems, different types of data transmission, and large number of tables, the configuration is complex and time-consuming.

Method used

A data fast upload engine implementation method based on DataX is designed, including analytical conversion module, data reading module, field mapping configuration module, configuration file generation module, upload module and scheduling module. Through the intelligent matching algorithm and visual field mapping configuration interface, the field mapping between the data template and the database table is automatically completed, and the json configuration file of the DataX execution instance is dynamically generated.

Benefits of technology

It realizes convenient, efficient and fast data upload, simplifies the mapping relationship maintenance between data import templates and database tables, improves configuration efficiency and accuracy, reduces operation complexity and workload, and supports the rapid upload and import of Excel or csv format data files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115729938B_ABST
    Figure CN115729938B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for implementing a data rapid upload engine based on DataX, belonging to the field of big data technology. It stipulates data templates and database tables corresponding to data upload corresponding business functions; provides a visual configuration interface for field mapping and an intelligent matching algorithm based on regular expressions; implements a json configuration file generator for DataX execution; encapsulates an SDK and provides a unified standard API for integrated execution of file upload and data import. The absolute path of the final csv format data file is backfilled into the json configuration to generate a DataX executable json configuration file, and the DataX tool is driven in the way of java, shell or python to complete the data synchronization task, solving the technical problem of providing a general data upload program component to achieve convenient, efficient and rapid data upload, simplifying the development of data upload and import functions and improving the development efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data, and particularly relates to a method for implementing a data rapid upload engine based on DataX. Background Art

[0002] In various daily business application systems, Excel data import seems to be a very common basic function. Taking a project mainly based on JavaWeb technology as an example, it usually uses the technology of the Apache POI project to implement the parsing of Excel, writes relevant SQL statements, and combines with JDBC technology to implement data import. It often needs to complete the development and implementation according to the Excel data template and the structure of the data table for each function. On the one hand, it is necessary to carry out customized development for each import function, consuming a lot of development effort; on the other hand, due to the limited experience and technical level of ordinary programmers, most of the implemented import functions have low efficiency.

[0003] DataX is an offline synchronization tool for heterogeneous data sources, dedicated to realizing stable and efficient data synchronization functions between various heterogeneous data sources. DataX has the characteristics of transmitting data using the http protocol, having a built-in reconnection and retry mechanism, providing distributed transactions and data consistency protection, high-speed data exchange between heterogeneous databases / file systems, and completing all-memory operations within a single process during the transmission process. The DataX plugin system, as an ecosystem, provides a rich set of Reader plugins and Writer plugins for common data sources, and supports custom development of new plugins. However, the execution instance of DataX depends on a pre-compiled json configuration file (including the number of execution threads, read plugins and corresponding read data sources, tables and fields, write plugins and target write data sources, tables and fields). In actual use, there are often many fields in the json configuration file. In the application scenarios of multiple systems, different types of data transmission, and a large number of tables, the configuration is cumbersome, the efficiency is low, and the workload is very large. Therefore, when integrating DataX as a built-in general functional component into a specific project, it is urgent to solve the problem of dynamic generation of the json configuration file of the DataX execution instance. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for implementing a data rapid upload engine based on DataX, which solves the technical problem of providing a general data upload program component to achieve convenient, efficient, and rapid data upload.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A method for implementing a data rapid upload engine based on DataX includes the following steps:

[0006] Step 1: Build a parsing and transformation module, a data reading module, a field mapping configuration module, a configuration file generation module, an upload module, and a scheduling module in a distributed server;

[0007] Step 2: The parsing and transformation module presets a convention template. According to the convention template, it receives an imported Excel or csv format file, that is, a template file. Through the Apache POI project technology, it parses the template file to obtain the data column names and corresponding column index information in the title row of the template file;

[0008] Step 3: The data reading module uses jdbc or odbc technology to connect to the database and reads a detailed list of field lists from the database system table according to the data table name;

[0009] Step 4: The field mapping configuration module is responsible for writing a field intelligent matching algorithm based on regular expression technology to match the column names in the title row of the data template with the field names in the database table and establish a mapping relationship;

[0010] Step 5: The field mapping configuration module creates a visual field mapping configuration interface for the template file to display the automatically mapped field relationships through the intelligent matching algorithm;

[0011] Step 6: The configuration file generation module provides a json configuration file generator for a DataX execution instance, and dynamically generates a json configuration file for DataX execution according to the field mapping configuration of the template file;

[0012] Step 7: The upload module provides a general data upload and import SDK to complete the file upload;

[0013] Step 8: The upload module provides XlsToCsv and XlsxToCsv utility classes to convert xls and xlsx Excel format data files into csv format to support the upload and import of Excel format data files;

[0014] Step 9: The scheduling module directly calls DataX using a method encapsulated in Java, or uses shell or python to trigger the execution of the DataX data synchronization task through the python script provided by the DataX tool.

[0015] Preferably, when executing Step 2, the imported Excel or csv format files both adopt the preset convention template. The first row of the convention template is the title row, and the second row and onwards are data rows.

[0016] Preferably, when executing Step 4, the steps of establishing the mapping relationship specifically include:

[0017] Step S4-1: First, fully match the column names in the data template header row with the field names in the database table, and establish a mapping relationship for the matched ones.

[0018] Step S4-2: Then, fully match the unmatched template column names with the remarks of the remaining fields in the database table, and establish a mapping relationship.

[0019] Step S4-3: For the remaining unmatched template column names, perform a similarity match with the remarks of the remaining fields in the database table according to the similarity set by the rules.

[0020] Preferably, when performing Step 5, the visual field mapping configuration interface supports manual adjustment of the mapping relationship.

[0021] Preferably, when performing Step 5, the visual field mapping configuration interface is switched and displayed in the form of a tab page for each database table. The tab page header shows the database table name and table remarks, and the following is displayed in tabular form.

[0022] Preferably, when performing Step 6, a json configuration file for DataX execution is automatically generated according to the preset rules and field mapping configuration, specifically including the Reader plugin part, the Writer plugin part, and the Transfer module part for data cleaning and transformation.

[0023] Preferably, when performing Step 6, when actually executing the DataX task, the absolute path of the csv format data file needs to be filled back into the json to generate an executable json file.

[0024] Preferably, when performing Step 7, the general data upload and import SDK specifically includes the upload of integrated files, the format verification of data files, the matching verification with data templates, the generation of the json configuration file for the DataX executable instance, and the invocation of the DataX tool.

[0025] A method for implementing a data rapid upload engine based on DataX according to the present invention solves the technical problem of providing a general data upload program component to achieve convenient, efficient, and rapid data upload, provides a visual field mapping component to implement the mapping relationship between the data import template (Excel or csv format) for each business function and the database table, and the general maintenance configuration function of the mapping relationship between the columns of the import template and the fields of the database table. By writing an intelligent matching algorithm, the intelligent automatic matching and mapping between the column names in the title row of the template and the field names or field remarks in the database table are realized, greatly reducing the operation complexity and workload of users, and improving the configuration efficiency and configuration accuracy. A json configuration file generator for a dedicated DataX execution instance is realized: provides a method for generating a json configuration file required for DataX execution according to the field mapping configuration of the data import template, the data source of the application system, and the database table, and the address of the temporarily missing actual csv format data file (the path attribute value in the txtfilereader plugin is a specific placeholder); and a method for filling the address of the actual specific csv format data file back into the json configuration file. Thus, the dynamic generation ability of the json configuration file of the DataX execution instance is realized. It well solves the problem that in the actual application of DataX, there are many field columns in the json configuration file, and in the application scenarios of multiple systems, different types of data transmission, and a large number of tables, the configuration is cumbersome and the efficiency is low. It provides verification on whether the uploaded data file matches the format specification of the function import template, integrates the ability to convert Excel to csv, and the execution scheduling strategy of DataX. Thus, the ability to rapidly upload and import data files in Excel or csv format is realized.

[0026] The present invention greatly simplifies the development of the data upload and import function in the application system, saves the R & D cost, improves the development efficiency, and ensures the reliability of business execution and the writing efficiency of a large amount of data. At the same time, the present invention is also very suitable for business scenarios involving a large amount of data interaction based on Excel or csv format files. Brief Description of the Drawings

[0027] Figure 1 is the flow chart of the present invention;

[0028] Figure 2 is the logic diagram of the method for implementing the data rapid upload engine based on DataX of the present invention. Detailed Embodiments

[0029] As Figure 1 - Figure 2 shown, a method for implementing a data rapid upload engine based on DataX includes the following steps:

[0030] Step 1: Build a parsing and transformation module, a data reading module, a field mapping configuration module, a configuration file generation module, an upload module, and a scheduling module in a distributed server;

[0031] Step 2: The parsing and transformation module presets a convention template. According to the convention template, it receives the imported Excel or csv format file, that is, the template file, and parses the template file through the Apache POI project technology to obtain the data column names and corresponding column index information in the title row of the template file;

[0032] Both the imported Excel or csv format files adopt the preset convention template. The first row of the convention template is the title row, and the second row is the data row.

[0033] The present invention maintains a template for the business function that needs to upload and import data and the corresponding database table to which the data is to be written; and it is stipulated that the first row of the template is the title row, and the second row and onwards are the data rows.

[0034] The present invention uses technologies such as Apache POI to implement the parsing of the title row of the Excel or csv format template to obtain information such as the title column names and corresponding column indexes.

[0035] In this embodiment, it is necessary to limit the format of the data import template to Excel (including the formats with xls and xlsx suffixes) or csv, and it is agreed that the first row of the template is the title row, and the second row and onwards are the data rows. By introducing project technologies such as Apache POI, the template file in Excel or csv format is parsed to obtain the data column names and corresponding column index information in the title row.

[0036] Step 3: The data reading module connects to the database using jdbc or odbc technology and reads a detailed list of field lists from the database system table according to the data table name; in this embodiment, the list of field lists includes information such as field names, field types, and field remarks.

[0037] The user of the current data source must have the query permission for the database system table in order to obtain the list of field information in the relevant data tables.

[0038] Step 4: The field mapping configuration module is responsible for writing a field intelligent matching algorithm implemented based on regular expression technology to match the column names in the title row of the data template with the field names in the database table and establish a mapping relationship;

[0039] When performing Step 4, the steps of establishing the mapping relationship, that is, the regular expression technology, for writing the intelligent matching algorithm specifically include:

[0040] Step S4-1: First, fully match the column names in the data template header row with the field names in the database table, and establish a mapping relationship for the matched ones;

[0041] Step S4-2: Then, fully match the unmatched template column names with the remarks of the remaining fields in the database table to establish a mapping relationship;

[0042] Step S4-3: For the remaining unmatched template column names, perform a similarity match with the remarks of the remaining fields in the database table according to the similarity set by the rules.

[0043] The present invention performs a rule-based match between the column names in the template header row and the field names or field remarks in the database table, and automatically establishes a mapping relationship; first, pair them according to a complete string match, and then pair the database table fields that are not paired by the complete match according to the string match similarity. If the data import template corresponds to multiple database tables, multiple mapping relationships will be intelligently generated.

[0044] Step 5: The field mapping configuration module creates a visual field mapping configuration interface for the template file to display the automatically mapped field relationships through the intelligent matching algorithm;

[0045] When performing Step 5, the visual field mapping configuration interface supports manual adjustment of the mapping relationship.

[0046] The visual field mapping configuration interface is switched and displayed in the form of a tab page for each database table. The tab page header shows the database table name and table remarks, and the following is displayed in a table form. Specifically, in this embodiment, when the template corresponds to multiple database tables, it is switched and displayed in the form of a tab page for each database table; the tab page header shows the database table name and table remarks, and the following is displayed in a table form as follows: data import template column index, column name in the data import template header row, database table field name (drop-down box, and the mapped and matched field is automatically selected, and the drop-down option range is the currently mapped field and all the fields in this database table that have not been mapped and matched), database table field type (linked with the field name column), database table field remarks (linked with the field name column), enumeration value conversion relationship (key-value key-value pair list, each key is separated by an English comma, for example, "1-male,0-female"); the database table field name column provides a drop-down editing function to support manual adjustment of the mapping relationship between the data import template column and the database table field.

[0047] Step 6: The configuration file generation module provides a json configuration file generator for a DataX execution instance, and dynamically generates a json configuration file for DataX execution according to the field mapping configuration of the template file;

[0048] When performing step 6, the generated json configuration file used by DataX execution specifically includes the Reader plugin part, the Writer plugin part, and the Transfer module part for data cleaning and transformation.

[0049] When saving the field mapping configuration of the data import template, the content of the json configuration file of DataX is generated. However, at this time, the json still lacks the actual execution data file address. When actually executing the DataX task, the absolute path of the csv format data file needs to be filled back into the json to generate an executable json file.

[0050] In this embodiment, a basic template of a DataX execution instance json configuration file is specifically defined to implement the basic framework structure of json, including the DataX execution speed configuration, etc. The main content part (including the Reader plugin part, the Writer plugin part, the Transformer data conversion module part, etc.) is temporarily replaced by multiple corresponding placeholders.

[0051] The Reader plugin part (mainly the txtfilereader plugin, the data file address and the column information of the corresponding data template), the Writer plugin part (the target library data source connection information, tables and fields), and the possible Transfer module for data cleaning and transformation (custom groovy functions to implement the conversion function of enumeration values in mapping matching).

[0052] The specific steps in this embodiment are as follows:

[0053] Step S6-1: Implement a method for generating the json text of the Reader plugin part. According to the field mapping configuration of the data import template, use the txtfilereader plugin provided in the DataX standard plugin library that supports csv reading to generate the json text of the Reader plugin part (configure to skip the first row header row of the data file through the "skipHeader": "true" attribute). At this time, the value of the path attribute in the txtfilereader plugin is still replaced by a specific placeholder, and when actually uploading, it can be replaced with the real file absolute path.

[0054] Step S6-2: Implement a method for generating partial JSON text of the Writer plugin. Generate the corresponding partial JSON text of the Writer plugin (including data source information, database table names, field names, etc.) according to the field mapping configuration of the data import template and the reading order of the template columns in the Reader plugin part. Use the corresponding Writer plugin provided by the standard plugin library of DataX according to the database type of the target data source. The data source can provide a method to directly read from the data source configuration file of the business application system using the present invention.

[0055] Step S6-3: Implement a method for generating partial JSON text of the Transformer data conversion module. If there is configuration data of "enumeration value conversion relationship (key-value pair list)" in the field mapping configuration of the data import template, generate the text content of the transformer field conversion module based on the custom groovy function according to the rules of the DataX execution instance JSON configuration file; if there is no configuration of "enumeration value conversion relationship", an empty string can be returned.

[0056] Step S6-4: Backfill and replace the partial JSON texts of the Reader plugin part, Writer plugin part, and Transformer data conversion module part into the basic template of the predefined DataX execution instance JSON configuration file, and return and save it.

[0057] Step 7: The upload module provides a general data upload and import SDK to complete the file upload;

[0058] To support the upload and import of data files in Excel format, it is necessary to implement the function of converting data files in xls and xlsx Excel formats into csv format.

[0059] Since the txtfilereader plugin provided in the DataX standard plugin library is used to read data in the present invention, the data files in Excel format must be converted into csv format to correctly read the data.

[0060] When executing Step 7, the general data upload and import SDK specifically includes the integration of file upload, format verification of data files, verification of matching with data templates, generation of the JSON configuration file of the DataX executable instance, and invocation of the DataX tool.

[0061] The present invention provides a general data upload and import SDK. It completes the verification of whether the uploaded data file matches the format specification of the function import template, integrates the ability to convert Excel to csv, and the execution scheduling strategy for DataX.

[0062] Step 8: The upload module provides XlsToCsv and XlsxToCsv utility classes to convert data files in two Excel formats, xls and xlsx, into csv format to support the upload and import of data files in Excel format.

[0063] In this embodiment, a standard and normative SDK is implemented using a specific programming language (such as Java, but not limited to Java), and an open API method for unified invocation and execution of file upload and data import is provided.

[0064] The specific steps in this embodiment are as follows:

[0065] Step S8-1: Provide a general file upload method and verify whether the file is a legal Excel or csv format file.

[0066] Step S8-2: Inside the SDK, implement the verification function of the uploaded data file and the data import template with the configured mapping rules, and verify whether the uploaded data file matches the format specification of the template and whether it is consistent with the data columns in the template (that is, whether the column names and orders of the first row and the first row title row of the template are both the same, but it is allowed to have a few more columns than in the template).

[0067] Step S8-3: Implement a tool package for converting Excel to csv in the SDK based on the Apache POI project technology, including XlsToCsv and XlsxToCsv utility classes; they are respectively used to convert the data files with xls suffix or xlsx suffix uploaded by users into csv format data files to suit the data reading of the txtfilereader plugin of DataX.

[0068] After the data file is uploaded, the non-csv format file is first converted into a csv format file, and the method provided by the json configuration file generator of the DataX execution instance is called to backfill the absolute path of the csv format data file into the corresponding json configuration file of the data import template, so as to generate the final json configuration file that can be used for DataX execution.

[0069] Step S8-5: Specify the generated instantiated json configuration file, schedule the DataX tool, and complete the task of reading data from the csv format data file, cleaning and converting through the transfer module, and finally writing it into the target database table.

[0070] The present invention realizes the convenient and efficient configuration of visual online field mapping by writing an automatic field mapping algorithm; realizes a json configuration file generator for a dedicated DataX execution instance, provides a method for generating a json configuration file of DataX for a specific data import template, and a method for backfilling the absolute path of a csv format file for the actually uploaded data file to generate a json configuration file of a DataX execution instance; provides a general API that includes a standard implementation for integrated call execution of file upload and data import, and an SDK toolkit for converting Excel files to csv files. Thus, a data rapid upload engine based on DataX that supports data file formats of Excel (with xls or xlsx suffix) or csv is realized.

[0071] Step 9: The scheduling module directly calls DataX by using a method encapsulated in Java, or uses the shell or python method to trigger the execution of the DataX data synchronization task through the python script provided by the DataX tool.

[0072] It is necessary to backfill the absolute path of the csv format data file to generate a json configuration file executable by DataX for DataX to call. Since DataX is implemented based on the Java language, the corresponding method can be encapsulated in Java to directly call DataX. It is also possible to use the shell or python method to trigger the execution of the DataX data synchronization task through the python script provided by the DataX tool.

[0073] Such as Figure 2The following is a logic diagram of the implementation method of the data rapid upload engine based on DataX of the present invention, which describes the whole process of using the data rapid upload engine based on DataX for data upload and import. First, for the business function that needs to implement data file upload and import, maintain the corresponding data upload template and related database tables; the background intelligent matching algorithm automatically completes the mapping of the data template columns and the fields in the database table, but there may be some cases where some columns and fields cannot be automatically matched; the successfully mapped situation and the template columns that have not been mapped can be viewed on the field mapping visualization configuration interface, and the mapping relationship between the template columns and the fields can be manually established or adjusted; when saving the field mapping relationship, the corresponding json configuration content of the data template and the database table is automatically assembled according to the requirements and specifications of the DataX tool; the business function user uploads a data file in Excel or csv format, and the engine automatically performs format verification on the data file and verifies whether it matches the corresponding data template; for the data file that passes the verification, if it is not in csv format, the Excel format data file needs to be converted into csv format through the encapsulated SDK tool class; finally, using the absolute path of the csv format data file, fill back the json configuration of the corresponding data template and the database table to generate the final executable json configuration file, and drive the DataX tool to execute the data synchronization task of the csv format data file to the target database table.

[0074] A method for implementing a data rapid upload engine based on DataX according to the present invention solves the technical problem of providing a general data upload program component to achieve convenient, efficient, and rapid data upload. It provides a visual field mapping component to achieve the mapping relationship between the data import template (Excel or csv format) for each business function and the database table, as well as the general maintenance configuration function for the mapping relationship between the columns of the import template and the fields of the database table. By writing an intelligent matching algorithm, it realizes the intelligent automatic matching and mapping between the column names in the title row of the template and the field names or field remarks in the database table, greatly reducing the operation complexity and workload of users, and improving the configuration efficiency and configuration accuracy. It realizes a json configuration file generator for a dedicated DataX execution instance: provides a method for generating a json configuration file required for DataX execution according to the field mapping configuration of the data import template, the data source of the application system, and the database table, where the address of the temporarily missing actual csv format data file (the value of the path attribute in the txtfilereader plugin is a specific placeholder); and a method for filling back the address of the actual specific csv format data file into the json configuration file. Thus, it realizes the dynamic generation ability of the json configuration file of the DataX execution instance. It well solves the problem that in the actual application of DataX, there are many field columns in the json configuration file, and in the application scenarios of multiple systems, different types of data transmission, and a large number of tables, the configuration is cumbersome and the efficiency is low. It provides a verification of whether the uploaded data file matches the format specification of the function import template, integrates the ability to convert Excel to csv, and the execution scheduling strategy of DataX. Thus, it realizes the ability to rapidly upload and import data files in Excel or csv format.

[0075] The present invention greatly simplifies the development of the data upload and import function in the application system, saves the R & D cost, improves the development efficiency, and ensures the reliability of business execution and the writing efficiency of a large amount of data. At the same time, the present invention is also very suitable for business scenarios involving a large amount of data interaction based on Excel or csv format files.

Claims

1. A method for implementing a data rapid upload engine based on DataX, characterized in that: It includes the following steps: Step 1: Build a parsing and transformation module, a data reading module, a field mapping configuration module, a configuration file generation module, an upload module, and a scheduling module in a distributed server; Step 2: The parsing and transformation module presets a convention template, receives an imported Excel or csv format file, i.e., a template file, according to the convention template, and parses the template file through the Apache POI project technology to obtain the data column names and corresponding column index information in the title row of the template file; Step 3: The data reading module connects to the database using jdbc or odbc technology and reads a detailed list of field lists from the database system table according to the data table name; Step 4: The field mapping configuration module is responsible for writing a field intelligent matching algorithm based on regular expression technology to match the column names in the data template title row with the field names in the database table and establish a mapping relationship; Step 5: The field mapping configuration module creates a visual field mapping configuration interface for the template file to display the automatically mapped field relationships through the intelligent matching algorithm; Step 6: The configuration file generation module provides a json configuration file generator for a DataX execution instance, and dynamically generates a json configuration file for DataX execution according to the field mapping configuration of the template file; Step 7: The upload module provides a general data upload and import SDK to complete the file upload; Step 8: The upload module provides XlsToCsv and XlsxToCsv utility classes to convert xls and xlsx Excel format data files into csv format to support the upload and import of Excel format data files; Step 9: The scheduling module directly calls DataX using a method encapsulated in Java, or uses the shell or python method to trigger the execution of the DataX data synchronization task through the python script provided by the DataX tool.

2. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When executing Step 2, the imported Excel or csv format files all adopt the preset convention template. The first row of the convention template is the title row, and the second row and onwards are data rows.

3. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When executing Step 4, the steps for establishing the mapping relationship specifically include: Step S4-1: First, completely match the column names in the data template title row with the field names in the database table, and establish a mapping relationship for those that match; Step S4-2: For the template column names that fail to match, then completely match them with the remarks of the remaining fields in the database table to establish a mapping relationship; Step S4-3: For the remaining unmatched template column names, then perform a similarity match with the remarks in the remaining fields of the database table according to the similarity set by the rules.

4. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When executing Step 5, the visual field mapping configuration interface supports manual adjustment of the mapping relationship.

5. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When executing Step 5, the visual field mapping configuration interface is switched and displayed in the form of a tab page for each database table. The tab page header shows the database table name and table remarks, and the following is displayed in tabular form.

6. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When performing step 6, the generated json configuration file used by DataX execution specifically includes the Reader plugin part, the Writer plugin part, and the Transfer module part for data cleaning and transformation.

7. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When performing step 6, it is necessary to fill back the absolute path of the csv format data file into the json when actually executing the DataX task to generate an executable json file.

8. The method for implementing a data rapid upload engine based on DataX according to claim 1, characterized in that: When performing step 7, the general data upload and import SDK specifically includes the upload of integrated files, the format verification of data files, the matching verification with data templates, the generation of the json configuration file for the executable instance of DataX, and the invocation of the DataX tool.

Citation Information

Patent Citations

  • Data synchronization method, device and system, computing device and storage medium

    CN108121757A

  • Data transmission method and device based on data x

    CN114490892A