Method and device for writing heterogeneous data into star ring ArgoDB based on DataX

Through DataX's read plug-in and custom ArgoDB write plug-in, efficient writing of heterogeneous data sources to starring ArgoDB is achieved, solving the problems of inefficiency and insufficient scalability when synchronizing heterogeneous data sources to ArgoDB, and achieving efficient and stable data synchronization effect.

CN120162005APending Publication Date: 2025-06-17CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510294432.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art has problems such as inefficiency, limited data source support, insufficient scalability and poor write performance when synchronizing heterogeneous data source data to Starring ArgoDB.

Method used

Through DataX's read plug-in and custom ArgoDB write plug-in, efficient writing of heterogeneous data sources is achieved. The specific steps include reading data, writing to HDFS files in parallel, creating external surfaces and Holodesk tables in ArgoDB, and completing the final storage and integration of the data.

Benefits of technology

It realizes efficient data synchronization between heterogeneous data sources and starring ArgoDB, solves problems such as inefficiency and insufficient scalability in traditional methods, and improves the performance and reliability of data synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162005A_ABST
    Figure CN120162005A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for writing heterogeneous data into a star ring ArgoDB based on DataX. The method comprises the following steps: reading data from a source end heterogeneous data source by utilizing a reading plug-in of the DataX, and transmitting the data to a Channel channel of the DataX; acquiring data from the Channel channel through a self-defined ArgoDB write plug-in, and writing the data into an HDFS (Hadoop Distributed File System); creating an appearance in the ArgoDB, and associating the HDFS file with the appearance; and creating a Holodesk table in the ArgoDB, and importing the data in the appearance into the Holodesk table to complete the storage and integration of the data. The method is compatible with an existing DataX ecological system, source end data reading logic does not need to be greatly modified, and the implementation cost is reduced; parallel reading and writing are supported, and the data synchronization efficiency is remarkably improved; the high throughput and distributed storage capability of the HDFS are utilized, so that the resource utilization rate is improved; and through combination of the appearance and the Holodesk table, flexible storage and efficient query of the data are realized, and support is provided for large-scale data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of big data reading and writing, and particularly relates to a method, device, computer-readable storage medium, and electronic device for writing heterogeneous data into StarRocks ArgoDB based on DataX. Background Art

[0002] With the rapid development of big data technology, enterprises are facing an urgent need to extract data from various heterogeneous data sources (including relational databases, NoSQL databases, log files, etc.) and efficiently write it into a distributed database for storage and analysis. As a high-performance distributed database, StarRocks ArgoDB has the ability to store and analyze massive amounts of data. However, in practical applications, how to efficiently synchronize the data of heterogeneous data sources to ArgoDB remains a huge challenge for technical personnel.

[0003] Traditional data synchronization methods have the following problems: (1) Low efficiency: The data import tool provided by ArgoDB cannot achieve automated data import, requiring a large amount of additional work, resulting in low efficiency.

[0004] (2) Limited data source support: It is difficult to meet diverse data synchronization requirements, such as relational databases, NoSQL databases, and log files.

[0005] (3) Insufficient scalability: Existing open-source data synchronization tools or frameworks lack direct support for ArgoDB. When the data source type needs to be extended, a large amount of manual adaptation work is required.

[0006] (4) Poor writing performance: There are performance bottlenecks when inserting data using the traditional SQL INSERT INTO method.

[0007] Although DataX, an open-source data synchronization tool, supports the reading and writing of multiple data sources, it currently lacks direct support for the ArgoDB data source. Therefore, how to efficiently and stably synchronize heterogeneous data sources to ArgoDB based on DataX has become an urgent problem to be solved. Summary of the Invention

[0008] To address the above problems, this application proposes a new method for writing heterogeneous data into StarRocks ArgoDB based on DataX, aiming to solve the problem of efficiently and stably synchronizing heterogeneous data sources to the ArgoDB database.

[0009] This solution realizes the efficient writing of source-side data through the read plugin of DataX and a custom ArgoDB write plugin. The main implementation process is as follows: (1) Data reading: Use the read plugin of DataX to read data from the source database and transfer the data to the DataX Channel. This process supports parallel reading to improve the data reading efficiency of the source side.

[0010] (2) Data writing: Through a custom ArgoDB write plugin, obtain the source data from the Channel in parallel, and generate HDFS files with the data by means of the FileSystem API, while specifying the storage path of the HDFS files.

[0011] (3) Create an external table: Create an external table in ArgoDB and ensure that the file path of the external table is the same as the storage path of the HDFS files.

[0012] (4) Data import: Create a Holodesk table in ArgoDB, and import the data in the external table into the Holodesk table through SQL statements to complete the final storage and integration of the data.

[0013] Through the above method, this solution realizes the efficient data synchronization between heterogeneous data sources and StarRocks ArgoDB, and solves the problems such as low efficiency and insufficient scalability existing in traditional methods.

[0014] Generally speaking, the core design and advantages of this solution are as follows: (1) Seamless integration of DataX and ArgoDB: By developing a custom ArgoDB write plugin, seamless docking of DataX and ArgoDB is achieved, supporting the efficient synchronization of data from multiple heterogeneous data sources (such as relational databases, NoSQL databases, log files, etc.) to the ArgoDB database.

[0015] (2) HDFS as a data transfer station: Utilize the high throughput and distributed storage capabilities of HDFS, combined with the parallel reading and writing capabilities of DataX, to achieve efficient data transmission and storage, ensuring high performance and high reliability in the data synchronization process.

[0016] (3) Combination of external table and Holodesk table: By creating an external table and a Holodesk table, and using HDFS files as the data storage medium, flexible data storage and efficient query are achieved, taking into account both the flexibility of data synchronization and the optimization of query performance.

[0017] Specifically, to achieve the above objectives, this application provides the following technical solutions: The first aspect of this application provides a method for writing heterogeneous data into StarRocks ArgoDB based on DataX, and the method includes: S1. Use the DataX read plugin to read data from the source heterogeneous data source and transfer the data to the DataX Channel; S2. Obtain data from the Channel through a custom ArgoDB write plugin and write the data into the HDFS file system; S3. Create an external table in ArgoDB and associate the data file in the HDFS file system with the external table; S4. Create a Holodesk table in ArgoDB and import the data in the external table into the Holodesk table to complete the storage and integration of the data.

[0018] Furthermore, in the method of the present application, in step S1, it further includes that the DataX read plugin supports parallel reading. By configuring the split field and parallelism, the reading efficiency of the source data is improved.

[0019] Furthermore, in the method of the present application, in step S2, it further includes that the custom ArgoDB write plugin dynamically adjusts the parallel strategy of the write task according to whether the split field and parallelism are configured: When the split field and parallelism are configured, the write plugin splits the write logic into multiple subtasks and executes them in parallel; When the split field and parallelism are not configured, the write plugin uses a single subtask to execute the write logic.

[0020] Furthermore, in the method of the present application, in step S2, it further includes generating the data into an HDFS file through the FileSystem API and specifying the storage path, file type, and field delimiter of the file.

[0021] Furthermore, in the method of the present application, in step S3, it further includes that when creating the external table, the field attributes, delimiter, and file path of the external table are specified through SQL statements to ensure that the file path of the external table is consistent with the storage path of the HDFS file.

[0022] Furthermore, in the method of the present application, in step S4, it further includes importing the data in the external table into the Holodesk table through SQL statements to achieve the optimization of data storage and query.

[0023] The second aspect of the present application provides two devices for writing heterogeneous data based on DataX into StarRocks ArgoDB.

[0024] The first device includes: A data reading module that uses the DataX read plugin to read data from the source heterogeneous data source and transfer the data to the DataX Channel; The data writing module fetches data from the Channel and writes it into the HDFS file system through a custom ArgoDB writing plugin; The external table creation module is used to create an external table in ArgoDB and associate the data files in the HDFS file system with this external table; The data import module is used to create a Holodesk table in ArgoDB and import the data in the external table into the Holodesk table to complete the storage and integration of data.

[0025] The second device includes: The DataX engine: As the core control unit of this device, it is responsible for the scheduling and execution of data synchronization tasks; The DataX read plugin: It is used to read data from the source data source and supports multiple heterogeneous data sources (such as relational databases, NoSQL databases, and file systems, etc.); The custom ArgoDB writing plugin: It is used to write data into the HDFS distributed file system and generate an ArgoDB external table associated with the HDFS data file; The Channel: As an intermediate cache for data transmission, it is used to efficiently transmit data between the DataX read plugin and the custom ArgoDB writing plugin; The HDFS distributed file system: As a data transfer station, it is used to store temporary data files and provides high-throughput and highly fault-tolerant data storage capabilities; The ArgoDB database: It is used to store the final data and supports efficient data query and analysis through the Holodesk table.

[0026] When the above two devices are running, they both implement the steps of the aforementioned method for writing heterogeneous data into StarRocks ArgoDB based on DataX.

[0027] The third aspect of this application provides an electronic device, including: a memory and a processor; The memory: It is used to store computer programs; The processor: It is used to execute the computer program to implement the steps of the aforementioned method for writing heterogeneous data into StarRocks ArgoDB based on DataX.

[0028] The fourth aspect of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the aforementioned method for writing heterogeneous data into StarRocks ArgoDB based on DataX.

[0029] In summary, the method for writing heterogeneous data into StarRocks ArgoDB based on DataX proposed in this application has the following advantages: (1) High efficiency: By leveraging the high-performance data reading ability of DataX and combining it with the optimized storage and query performance of the Holodesk table in ArgoDB, efficient data query and efficient writing to the StarRing ArgoDB database are achieved.

[0030] (2) Compatibility: It is compatible with the existing DataX ecosystem, eliminating the need for substantial modification of the source-side data reading logic, thus reducing the implementation cost; when expanding the source-side data source types, only adaptation according to the logic of the DataX read plugin is required.

[0031] (3) Data consistency: A two-step writing mechanism of external tables and Holodesk tables is adopted to ensure the integrity and consistency of data during the import process.

[0032] (4) Automation and ease of maintenance: The entire data synchronization process can be automatically executed in conjunction with scheduling tools, enhancing the maintainability of the system.

[0033] (5) Resource optimization: Through the distributed storage and computing capabilities of HDFS and ArgoDB, resource utilization is optimized, effectively reducing the single-point performance bottleneck.

[0034] (6) Convenience: The data synchronization process is simplified. As a simple and feasible solution, it reduces the operation complexity and operation and maintenance costs of data synchronization.

[0035] Other features and advantages of this application will be elaborated in detail in the subsequent description, or can be understood by implementing the relevant technical solutions of this application. The objectives and other advantages of this application can be achieved through the technical features and means clearly pointed out in the description, claims, and drawings, and obtained through the implementation process of these technical contents. Brief Description of the Drawings

[0036] To more clearly elaborate the technical solutions of the embodiments of this application, the drawings involved in the description of the embodiments will be briefly introduced below. It should be noted that the drawings only show some embodiments of this application. For those skilled in the art, other relevant drawings can be deduced based on these drawings without creative efforts.

[0037] Figure 1 It is the overall design architecture diagram of the solution of this application.

[0038] Figure 2 It is the overall implementation flowchart of the method for writing heterogeneous data from this application to the StarRing ArgoDB based on DataX.

[0039] Figure 3 It is the processing flowchart of the write plugin in the solution of this application.

[0040] Figure 4 This is the composition structure diagram of the first DataX-based device for writing heterogeneous data into StarRing ArgoDB proposed in this application.

[0041] Figure 5 This is the composition structure diagram of the second DataX-based device for writing heterogeneous data into StarRing ArgoDB proposed in this application.

[0042] Figure 6 This is the schematic structural diagram of the electronic device provided by the embodiments of this application. Detailed implementation manners

[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer and more understandable, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be clear that the described embodiments are only some of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0044] In this article, the term "including" and any form of its deformation (such as "including", "including") are open expressions and should be understood as "including but not limited to", that is, the listed content is not an exhaustive list and may also include other content not explicitly mentioned. The term "based on" should be understood as "at least partially based on", that is, the referred basis or condition may not be the only factor and may also involve other relevant factors. The term "an embodiment" should be understood as "at least one embodiment", that is, the described embodiment is not the only possible implementation, and there may also be other similar embodiments.

[0045] In this application, when the terms "one" and "multiple" are used to modify relevant elements or features, their expressions are illustrative rather than restrictive. Unless otherwise clearly stated in the context, "one" should be understood as "at least one", and "multiple" should be understood as "at least two". Those skilled in the art should reasonably interpret these terms according to the semantic and logical relationships in the context to ensure that they cover the possibility of "one or more".

[0046] Figure 2 The following shows the overall implementation process of the DataX-based method for writing heterogeneous data into StarRing ArgoDB provided by this application, including the following steps: S1. Use the read plugin of DataX to read data from the source heterogeneous data source and transfer the data to the Channel channel of DataX; S2. Obtain data from the Channel through a custom ArgoDB write plugin and write the data into the HDFS file system; S3. Create an external table in ArgoDB and associate the data files in the HDFS file system with this external table; S4. Create a Holodesk table in ArgoDB and import the data in the external table into the Holodesk table to complete the storage and integration of the data.

[0047] To more clearly elaborate on the technical solution of this application, the following will be further described through embodiments of specific scenarios.

[0048] As Figure 1 shown, this solution uses the read plugin of DataX and a custom ArgoDB write plugin to achieve efficient writing of source-side data. The main implementation process is as follows: I. Reading of heterogeneous data: Use the read plugin of DataX to read data from the source-side database and transfer the data to the DataX Channel. This process supports parallel reading to improve the reading efficiency of source-side data.

[0049] In the link of reading heterogeneous data, the data reading logic of DataX is adopted. The specific implementation is as follows: (1) Reading configuration When configuring the read plugin of DataX, it is necessary to clearly specify the type of the source-side data source (such as MySQL, Oracle, file system, etc.) and related connection information. If concurrent reading of source-side data is required, it is also necessary to configure the split field (splitPk) of the data table and the corresponding parallelism. The split field is usually an auto-incrementing primary key field.

[0050] Table 1 Example data of source-side data id name address 10001 Zhang San Beijing 10002 Li Si Shenzhen Table 2 Configuration information of the read plugin Split field splitPk id Parallelism 3 (2) Reading operation Read data from the source-side data source through the read plugin of DataX and transfer the data to the Channel. The read plugin and the write plugin use the Channel as an intermediate cache for data transmission to achieve efficient data transfer.

[0051] Adopting the above heterogeneous data reading solution based on DataX has the following advantages: Compatibility: Compatible with the existing DataX ecosystem, without significantly modifying the source-side data reading logic, reducing the implementation cost.

[0052] High efficiency: Support concurrent reading. By configuring the split field and parallelism, the concurrent ability of data reading is significantly improved, and the overall efficiency of data synchronization is enhanced.

[0053] II. Data writing: Through a custom ArgoDB write plugin (the processing flow of the write plugin is as Figure 3 shown), the source data is obtained in parallel from the Channel, and the data is generated into an HDFS file with the help of the FileSystem API. At the same time, the storage path of the HDFS file is specified. The main operations are as follows: This solution uses HDFS (Hadoop Distributed File System) as the intermediate storage layer. HDFS has the characteristics of high fault tolerance and high throughput, and is suitable for scenarios with high concurrency and large data volume storage.

[0054] During the writing process, according to whether the split field (splitPk) and parallelism (parallelism) are configured in the read plugin, the write plugin will dynamically adjust the parallel strategy of the writing task: When the read plugin configures the split field and parallelism (assuming the parallelism is n, n>1), the write plugin will split the writing logic into n sub-tasks (task) and execute the writing logic in parallel.

[0055] When the read plugin does not configure the split field and parallelism, the write plugin only uses a single sub-task (task) to execute the writing logic.

[0056] (1) Obtain the source data from the Channel In each sub-task (task), first obtain the source data from the corresponding Channel.

[0057] (2) Generation of HDFS files In each sub-task, after obtaining the source data, the data is generated into a file in the HDFS distributed file system through the FileSystem API and stored in the specified HDFS path (such as / default.db / table_target). The generated file type is a text file (such as 00001.txt), and at the same time, it supports specifying a custom delimiter (such as ",").

[0058] In this solution, multiple sub-tasks (task) cooperate with each other to generate HDFS files in parallel, thus significantly improving the file generation efficiency.

[0059] III. Create an external table in ArgoDB: Create an external table in ArgoDB and ensure that the file path of this external table is the same as the storage path of the HDFS file.

[0060] An external table in ArgoDB is a special type of database object that defines a readable data source. It is not a table inside the database but metadata pointing to the data in an external data source. The external table provides a way to associate external data files with the database, allowing users to use this external data in SQL queries just like operating on ordinary tables.

[0061] In ArgoDB, to create an external table to map the data files on HDFS, the DDL statement for creating the table is as follows: create external table text_a (id int, name string) row format delimited fields terminated by ',' location ' / default.db / table_target'; Among them: Separator specification: Specify the field separator of the external table through the 'terminated by' keyword. The separator here needs to be the same as the field separator of the text file (.txt) generated on HDFS.

[0062] Path mapping: The path after the 'location' keyword needs to be the same as the HDFS storage path specified in the previous step (such as / default.db / table_target).

[0063] Field mapping: The field attributes of the external table (such as id and name) need to be the same as those of the source data table to ensure the correct association and query of data.

[0064] IV. Data import: Create a Holodesk table in ArgoDB and import the data in the external table into the Holodesk table through SQL statements to complete the final storage and integration of the data.

[0065] (1) Create a Holodesk table in ArgoDB The Holodesk table is an optimized storage structure in ArgoDB, designed specifically for data analysis, real-time query, and large-scale data processing scenarios, with significant performance advantages. Create a Holodesk table in ArgoDB to store the final data. The table creation statement is as follows: create table holo_a (id int, name string) stored as holodesk; (2)Import the external table data into the Holodesk table Use an SQL statement to import the data in the external table text_a into the Holodesk table holo_a to achieve the final storage and optimization of the data. The import statement is as follows: insert into holo_a select * from text_a; This operation efficiently migrates the data in the external table to the Holodesk table, making full use of the storage and query optimization features of the Holodesk table to support subsequent data analysis and processing.

[0066] Figure 4 and Figure 5 Two devices for writing heterogeneous data to StarRing ArgoDB based on DataX proposed in this application are shown below.

[0067] The first device includes: A data reading module that uses the DataX read plugin to read data from the source heterogeneous data source and transfers the data to the DataX Channel; A data writing module that obtains data from the Channel through a custom ArgoDB write plugin and writes the data into the HDFS file system; An external table creation module for creating an external table in ArgoDB and associating the data file in the HDFS file system with the external table; A data import module for creating a Holodesk table in ArgoDB and importing the data in the external table into the Holodesk table to complete the storage and integration of the data.

[0068] The second device includes: The DataX engine: As the core control unit of this device, it is responsible for scheduling and executing data synchronization tasks; The DataX read plugin: Used to read data from the source data source and supports multiple heterogeneous data sources (such as relational databases, NoSQL databases, and file systems, etc.); The custom ArgoDB write plugin: Used to write data into the HDFS distributed file system and generate an ArgoDB external table associated with the HDFS data file; The Channel: As an intermediate cache for data transmission, it is used to efficiently transmit data between the DataX read plugin and the custom ArgoDB write plugin; HDFS distributed file system: As a data transfer station, it is used to store temporary data files and provides high-throughput and highly fault-tolerant data storage capabilities; ArgoDB database: It is used to store the final data and supports efficient data query and analysis through Holodesk tables.

[0069] When the above two devices are running, they both implement the steps of the method for writing heterogeneous data into Star Ring ArgoDB based on DataX disclosed in this application.

[0070] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementations of the devices, methods, and computer program products according to various embodiments of the present application, including the system architecture, functions, and operations. In these figures, each block may represent a module, a program segment, or a part of the code, which contains one or more executable instructions for implementing the specified logical function. It should be noted that each block in the block diagram and / or flowchart, as well as combinations of these blocks, can be implemented by a dedicated hardware-based system to perform the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0071] As Figure 6 shown, an embodiment of the present application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing computer programs executable by the processor, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 runs the executable computer program to implement the steps of the above-mentioned method for writing heterogeneous data into Star Ring ArgoDB based on DataX.

[0072] It can be understood that in addition to the memory and the processor, this electronic device may also include an input device (such as a keyboard), an output device (such as a display), and other communication modules. These input devices, output devices, and other communication modules communicate with the processor through the I / O interface (i.e., the input / output interface).

[0073] The operations of the present application can be implemented by writing computer program code using one or more programming languages or combinations thereof. The programming languages include but are not limited to the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; Conventional procedural programming languages, such as the "C" language or similar programming languages.

[0074] The execution modes of the program code include but are not limited to: Fully executed on the user's computer; Part is executed on the user's computer and part is executed on a remote computer; Execute as an independent software package; Execute entirely on a remote computer or server.

[0075] In scenarios involving a remote computer, the remote computer can be connected to the user's computer through any type of network connection, and the network includes but is not limited to a local area network (LAN) or a wide area network (WAN). In addition, the remote computer can also be connected to an external computer through an Internet service provider, for example, by using the Internet for connection.

[0076] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute each step of the method for writing heterogeneous data based on DataX into StarRing ArgoDB disclosed in the present application.

[0077] In the context of the present application, a computer-readable storage medium refers to a tangible medium that can store computer program code and related data. Specific examples include but are not limited to the following: (1) Portable computer disk: A removable magnetic storage medium such as a floppy disk.

[0078] (2) Hard disk: Fixed storage devices including mechanical hard disks and solid-state drives.

[0079] (3) Random access memory (RAM): A volatile storage medium for temporarily storing data and program code.

[0080] (4) Read-only memory (ROM): A non-volatile storage medium for storing fixed programs and data.

[0081] (5) Erasable programmable read-only memory (EPROM) or flash memory: A non-volatile storage medium that supports multiple erasures and programming.

[0082] (6) Fiber optic storage device: A storage medium based on fiber optic technology.

[0083] (7) Portable compact disc read-only memory (CD-ROM): A read-only medium for storing data in the form of an optical disc.

[0084] (8) Optical storage device: Storage media based on optical principles such as DVDs and Blu-ray discs.

[0085] (9) Magnetic storage device: Storage media based on magnetic principles such as magnetic tapes and disks.

[0086] (10) Any suitable combination of the above: for example, multiple storage media are combined for use to meet different storage requirements.

[0087] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.

[0088] In particular, according to the embodiments of this application, the processes described in the flowcharts can be implemented as computer software programs. For example, the embodiments of this application relate to a computer program product that includes a computer program carried on a non-transitory computer-readable medium. The computer program contains program code for executing the method for writing heterogeneous data based on DataX into Starring ArgoDB disclosed in this application. When the computer program is executed by a processing device, the above functions defined in the embodiments of this application can be realized.

[0089] Although there are several specific implementation details in the above discussion, these details should not be construed as limiting the scope of this application. The above description is only a preferred embodiment of this application and an explanation of the applied technical principles. Those skilled in the art should understand that the disclosed scope of this application is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept.

[0090] Those skilled in the art should also understand that they can modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features without departing from the spirit and scope of the technical solutions of the embodiments of this application. These modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for writing heterogeneous data into StarRing ArgoDB based on DataX, characterized in that: The method comprises: S1. Use DataX's read plug-in to read data from the source heterogeneous data source and transfer the data to the DataX Channel; S2. Get data from the Channel channel through the custom ArgoDB write plug-in and write the data to the HDFS file system; S3. Create a foreign table in ArgoDB and associate the data files in the HDFS file system with the foreign table. S4. Create a Holodesk table in ArgoDB and import the data in the foreign table into the Holodesk table to complete data storage and integration.

2. The method according to claim 1, characterized in that Step S1 also includes that the DataX read plug-in supports parallel reading, which improves the reading efficiency of source data by configuring segmentation fields and parallelism.

3. The method according to claim 1, characterized in that Step S2 also includes that the custom ArgoDB write plug-in dynamically adjusts the parallel strategy of the write task according to whether the segmentation field and parallelism are configured: When the segmentation field and parallelism are configured, the write plug-in divides the write logic into multiple subtasks and executes them in parallel; When the split field and parallelism are not configured, the write plugin uses a single subtask to execute the write logic.

4. The method according to claim 1, characterized in that Step S2 also includes generating the data into an HDFS file through the FileSystem API, and specifying the storage path, file type, and field separator of the file.

5. The method according to claim 4, characterized in that Step S3 also includes, when creating a table, specifying the field attributes, delimiters and file path of the table through an SQL statement to ensure that the file path of the table is consistent with the storage path of the HDFS file.

6. The method according to claim 1, characterized in that Step S4 also includes importing the data in the external table into the Holodesk table through SQL statements to achieve data storage and query optimization.

7. A device for writing heterogeneous data into StarRing ArgoDB based on DataX, characterized in that: The device comprises: The data reading module uses the DataX reading plug-in to read data from the source heterogeneous data source and transmit the data to the DataX Channel; The data writing module obtains data from the Channel channel through a custom ArgoDB writing plug-in and writes the data to the HDFS file system; The table creation module is used to create a table in ArgoDB and associate the data files in the HDFS file system with the table; The data import module is used to create a Holodesk table in ArgoDB and import data from the external table into the Holodesk table to complete data storage and integration.

8. A device for writing heterogeneous data into StarRing ArgoDB based on DataX, characterized in that: The device comprises the following components: DataX Engine: As the core control unit of this device, it is responsible for scheduling and executing data synchronization tasks; DataX read plug-in: used to read data from the source data source, supporting multiple heterogeneous data sources; Custom ArgoDB write plug-in: used to write data to the HDFS distributed file system and generate an ArgoDB external table associated with the HDFS data file; Channel: As an intermediate cache for data transmission, it is used to efficiently transmit data between the DataX read plug-in and the custom ArgoDB write plug-in; HDFS distributed file system: As a data transfer station, it is used to store temporary data files and provides high throughput and high fault tolerance data storage capabilities; ArgoDB database: used to store final data and support efficient data query and analysis through Holodesk tables.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for writing heterogeneous data based on DataX into StarRing ArgoDB as described in any one of claims 1 to 6 are implemented.

10. An electronic device, characterized in that: include: Memory and processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the method for writing heterogeneous data into StarRing ArgoDB based on DataX as described in any one of claims 1-6.

Citation Information

Cited By

  • Method and device for optimizing write-in Hive Parquet storage format based on DataX

    CN121524154A

  • A method and apparatus for optimizing writing to HiveParquet storage format based on DataX

    CN121524154B