Data extraction method and device, storage medium, equipment and program product

By adopting the methods of shard extraction and asynchronous upload in the business intelligence system and optimizing the data storage structure in the distributed column database, the performance bottleneck of traditional data extraction methods in massive data processing is solved, and efficient and accurate data extraction and query performance improvement is achieved.

CN119938729APending Publication Date: 2025-05-06HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411719174.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional data extraction methods have performance bottlenecks when facing massive data, such as slow extraction speed, high resource consumption, poor configuration flexibility, and difficult to meet the real-time, accuracy and efficiency requirements of business intelligence systems.

Method used

By multi-level splitting and asynchronous batch upload of shard extraction results, target tables and temporary tables are established in the distributed column database management system, and the data storage structure is optimized by fine extraction configuration parameters, improving extraction performance and improving query efficiency.

Benefits of technology

It significantly improves data processing efficiency and system resource utilization, realizes efficient, accurate and flexible data extraction operations, and meets users' needs for real-time data analysis and rapid report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938729A_ABST
    Figure CN119938729A_ABST
Patent Text Reader

Abstract

The invention discloses a data extraction method and device, a storage medium, equipment and a program product, and the method comprises the steps: receiving a data extraction task, and obtaining an extraction configuration parameter corresponding to the data extraction task; according to the extraction configuration parameters, establishing a target table and a temporary table corresponding to the data extraction task in a distributed column database management system; segmenting the data extraction task into a plurality of fragmentation tasks according to the extraction configuration parameters; executing each fragmentation task, and writing fragmentation extraction result data corresponding to each fragmentation task into a plurality of temporary files in batches; compressing a target temporary file of which the file size reaches a preset threshold in the plurality of temporary files, and asynchronously uploading the compressed target temporary file to a temporary table; and after all the fragmentation tasks are executed, the data in the temporary table is submitted to the target table, data extraction is completed, and the data processing efficiency and the system resource utilization rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data extraction method, device, storage medium, equipment and program product. Background Art

[0002] In the current Business Intelligence (BI) system, data extraction is a crucial link, which is directly related to the real-time, accuracy and efficiency of data analysis. Traditional data extraction methods often have performance bottlenecks when facing massive data, such as slow extraction speed, high resource consumption, poor configuration flexibility, etc. These problems seriously restrict the application effect and user experience of the BI system.

[0003] In particular, traditional data extraction methods are often unable to cope with scenarios that require processing large amounts of data. For example, in the process of data integration involving multiple data sources, the heterogeneity between different data sources increases the complexity of data extraction; at the same time, the processing of large amounts of data places higher demands on the system's concurrency and response speed.

[0004] In addition, current data extraction methods often have problems such as low efficiency and unreasonable data organization, resulting in poor query performance of BI reports. Many BI systems lack effective organization and optimization of data after data extraction, resulting in poor query performance and unable to meet users' needs for real-time data analysis and fast report generation.

[0005] In order to solve the above problems, the market is in urgent need of an efficient, flexible and performance-optimized data extraction method. Summary of the invention

[0006] The embodiments of the present application provide a data extraction method, apparatus, storage medium, device and program product, which can improve the extraction performance by multi-level splitting of the sharded extraction result data and asynchronous batch uploading, optimize the data storage structure through fine extraction configuration parameters, and improve query efficiency.

[0007] On the one hand, an embodiment of the present application provides a data extraction method, the method comprising:

[0008] Receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task;

[0009] According to the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in a distributed column database management system;

[0010] According to the extraction configuration parameters, the data extraction task is divided into a plurality of sharding tasks;

[0011] Execute each sharding task and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches;

[0012] Compressing a target temporary file whose file size reaches a preset threshold among the multiple temporary files and asynchronously uploading it to the temporary table;

[0013] When all sharding tasks are completed, the data in the temporary table is submitted to the target table to complete the data extraction.

[0014] On the other hand, an embodiment of the present application provides a data extraction device, the device comprising:

[0015] An acquisition unit, used to receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task;

[0016] An establishing unit, used to establish a target table and a temporary table corresponding to the data extraction task in a distributed column database management system according to the extraction configuration parameters;

[0017] A splitting unit, used for splitting the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters;

[0018] An execution unit, used to execute each sharding task and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches;

[0019] A processing unit, configured to compress a target temporary file whose file size reaches a preset threshold among the multiple temporary files and asynchronously upload the target temporary file to the temporary table;

[0020] The submitting unit is used to submit the data in the temporary table to the target table to complete data extraction after all sharding tasks are executed.

[0021] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the data extraction method described in any of the above embodiments.

[0022] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the data extraction method described in any of the above embodiments by calling the computer program stored in the memory.

[0023] On the other hand, an embodiment of the present application provides a computer program product, including computer instructions, which, when executed by a processor, implement the data extraction method described in any of the above embodiments.

[0024] The embodiment of the present application receives a data extraction task and obtains the extraction configuration parameters corresponding to the data extraction task; according to the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in a distributed column database management system; according to the extraction configuration parameters, the data extraction task is divided into multiple sharding tasks; each sharding task is executed, and the sharding extraction result data corresponding to each sharding task is written in batches to multiple temporary files; the target temporary files whose file sizes reach a preset threshold in multiple temporary files are compressed and asynchronously uploaded to the temporary table; when all sharding tasks are executed, the data in the temporary table is submitted to the target table to complete the data extraction. The embodiment of the present application realizes efficient, accurate and flexible data extraction operations by adopting advanced technical means such as distributed processing, parallel execution of sharding tasks, and asynchronous uploading, which significantly improves data processing efficiency and system resource utilization. Among them, the extraction performance is improved by multi-level splitting of the sharding extraction result data and asynchronous batch uploading, and the data storage structure is optimized by fine extraction configuration parameters, thereby improving query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flowchart of a data extraction method provided in an embodiment of the present application.

[0026] Figure 2 A schematic diagram of a first application scenario of the data extraction method provided in an embodiment of the present application.

[0027] Figure 3 A schematic diagram of a second application scenario of the data extraction method provided in an embodiment of the present application.

[0028] Figure 4 A schematic diagram of a third application scenario of the data extraction method provided in an embodiment of the present application.

[0029] Figure 5 A schematic diagram of the structure of a data extraction device provided in an embodiment of the present application.

[0030] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application.

[0031] Figure 7 A schematic diagram of a storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0033] The embodiments of the present application provide a data extraction method, apparatus, storage medium, device and program product. Specifically, the data extraction method of the embodiments of the present application can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a smart phone, a tablet computer, a laptop computer, a smart TV, a smart speaker, a wearable smart device, a smart car terminal and other devices. The terminal can also include a client, which can be a video client, a browser client, an instant messaging client or a small program. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0034] The embodiments of the present application can be applied to various scenarios such as data processing and data extraction.

[0035] First, some nouns or terms that appear in the description of the embodiments of the present application are explained as follows:

[0036] Business Intelligence (BI): refers to the use of modern data warehouse technology, online analytical processing technology, data mining and data presentation technology to analyze data to realize business value.

[0037] Massively Parallel Processing (MPP) database: A type of database designed for large-scale data processing. It processes data in parallel on multiple servers to support complex analysis and queries, and can handle massive amounts of data and provide fast query results.

[0038] Distributed column-based database management system (ClickHouse): an open source distributed column-based database management system suitable for big data scenarios with complex analysis and instant query.

[0039] Data extraction: The process of extracting data from a data source. Common forms include full extraction and incremental extraction. Its function is to transfer data to an intermediate storage or target storage location for further processing or analysis.

[0040] Structured Query Language (SQL): is a standard language for querying and operating databases, widely used in relational database management systems. It has multiple functions such as querying, inserting, updating, deleting, defining and managing database objects.

[0041] JDBC driver: Java Database Connectivity (JDBC) is a standard application programming interface (API) used by Java programs to connect to and operate databases. It acts as a bridge between Java applications and databases, enabling Java code to perform SQL queries, record updates, and other operations.

[0042] It should be noted that the order of description of the following embodiments is not intended to limit the priority order of the embodiments.

[0043] Each embodiment of the present application provides a data extraction method, which can be executed by a terminal or a server, or by a terminal and a server together; the embodiments of the present application are described by taking the data extraction method executed by a server as an example.

[0044] See also Figures 1 to 4 , Figure 1 A flow chart of a data extraction method provided in an embodiment of the present application is shown in FIG. Figures 2 to 4 A schematic diagram of an application scenario of the data extraction method provided in an embodiment of the present application. The method may include the following steps:

[0045] Step 110: receiving a data extraction task and obtaining extraction configuration parameters corresponding to the data extraction task.

[0046] like Figure 2 As shown, after receiving the data extraction task from the upper business platform, the extraction begins. The source of the extracted data is provided by the data access interface, which can obtain data from various online analytical processing (OLAP) / online transaction processing (OLTP) data sources and text data sources. For example, the data source may include a distributed database (such as Hive), a relational database (such as MySQL, Oracle), a spreadsheet file (Excel), a comma-separated value file (CSV), etc. The extracted data may include at least one of structured data and semi-structured data.

[0047] For example, the data extraction task may include task information, such as the connection information of the source database, the name of the target table, the data table and field to be extracted, etc. The data extraction task can be received in a variety of ways, including but not limited to API interface calls, network (Web) interface submissions, command line tool executions, etc. These channels should provide user-friendly interfaces or commands so that users can submit tasks conveniently.

[0048] At the same time, the configuration parameters of the data extraction task are also obtained, including but not limited to the extracted shard fields, partition fields, calculated primary keys, and field non-null value identification information.

[0049] Extraction sharding field: A field used to split a data extraction task into multiple subtasks. The system splits the extracted data result set based on the extraction sharding field, that is, the data extraction task SQL can be split into several sub-SQLs through the extraction sharding field.

[0050] Partition field: The field used to create a partition table in ClickHouse. The system will create a target table based on the set partition field. The created target table will be a partitioned table, and ClickHouse will store data according to the partitions.

[0051] Computed primary key: a field used for sorting and indexing data in ClickHouse.

[0052] Field non-null value (Nullable) identification information: Use the JDBC interface to detect whether the column attributes of the table allow NULL values. Users can manually set important fields to non-nullable. For example, ClickHouse only allows the storage of NULL fields of the Nullable type, such as: Int64 type fields storing NULL will automatically become 0. This behavior is unacceptable in the BI system. Nullable (Int64) can store integer values ​​and NULL values. However, the Nullable type is not recommended in ClickHouse and has a negative impact on performance. Since the data in the source database is unknown, it is not clear in advance whether the column attributes contain NULL. The embodiment of the present application will first detect whether the column attributes of the table are not Nullable through the JDBC interface. Users can manually set important dimensions and measurement fields to not Nullable to assist in judgment.

[0053] In some embodiments, before receiving the data extraction task, the method further includes:

[0054] The extraction configuration parameters are configured, wherein the extraction configuration parameters include at least one of extraction of fragmentation fields, partition fields, calculation of primary keys, and field non-null value identification information.

[0055] For example, before receiving a data extraction task, the system or user needs to carefully configure the extraction configuration parameters. By properly configuring the extraction configuration parameters, the system can ensure the smooth and efficient execution of the data extraction task while meeting the user's specific needs and ensuring data quality.

[0056] Step 120: Establish a target table and a temporary table corresponding to the data extraction task in a distributed column-based database management system according to the extraction configuration parameters.

[0057] For example, before creating the target table and temporary table, the system first needs to determine the database connection information, including the Java Database Connection (JDBC) catalog, schema, table, etc., in order to locate the specific database and table. For example, if the task is to extract data from a MySQL database, the system needs to connect to the MySQL database through JDBC and locate the specific table based on the provided Catalog, Schema, and Table information. If data is extracted from Hive, the system needs to connect to Hive through JDBC and locate the specific table or view based on the provided information.

[0058] For some complex data extraction tasks, you can use custom SQL to determine the final result set. The system supports the use of custom SQL to flexibly define the logic and scope of data extraction.

[0059] After confirming the extraction configuration parameters, the system will create the target table and the temporary table corresponding to the target table based on the partition fields in the extraction configuration parameters. The partition fields are used to create partition tables in ClickHouse, which helps optimize data query and storage performance.

[0060] For example, both the target table and the temporary table are distributed tables, which means they can store and query data across multiple nodes. At the same time, a replicated merge tree (ReplicatedMergeTree) local table is also established on each node to ensure high availability and fault tolerance of data.

[0061] For example, if the task has been executed before, it means that the target table already exists, and there is no need to recreate the target table. If the target table does not exist, a new target table is created according to the extraction configuration parameters.

[0062] For example, whether the target table exists or not, a temporary table needs to be created. Temporary tables are used to store intermediate data generated during data extraction to avoid data inconsistency caused by directly operating the target table during the extraction process. The reason why data is currently extracted to temporary tables is that if the target table is directly operated during the extraction process, the user's report query data will continue to change during the extraction process, which is unacceptable. Therefore, the data needs to be written to the temporary table first, and then written to the target table after the data extraction is completed. This ensures that the user's report query data will not be affected during the data extraction process.

[0063] In some embodiments, establishing a target table and a temporary table corresponding to the data extraction task in a distributed column-based database management system according to the extraction configuration parameters includes:

[0064] According to the partition field in the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in the distributed column-based database management system, wherein the target table is configured with a first local table of each replica node in the distributed column-based database management system, and the temporary table is configured with a second local table of each replica node in the distributed column-based database management system, and each temporary file corresponds to a replica node in the distributed column-based database management system.

[0065] For example, the target table and temporary table are configured with local tables of each replica node in the distributed column database management system. For example, for a temporary table shadow_A of a data extraction task, it will correspond to the second local table shadow_A_local on each replica (ClickHouse) node; similarly, the target table target_A will also correspond to the first local table target_A_local on each node.

[0066] In the data extraction process, each temporary file corresponds to a replica node in the distributed column database management system. This design helps to achieve parallel processing and load balancing of data.

[0067] Step 130: divide the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters.

[0068] In some embodiments, dividing the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters includes: dividing the data extraction task into a plurality of sharding tasks according to an extraction sharding field in the extraction configuration parameters.

[0069] For example, the system divides the data extraction task into multiple sharding tasks according to the extraction sharding fields in the extraction configuration parameters. The sharding field type can be integer, floating point, date, datetime, string, etc. To better assist in data sharding, the system supports field type conversion. For example, even if the data type of a column is String, if its format is actually a date type, the system can identify this through field type conversion and perform sharding accordingly.

[0070] In some embodiments, the data extraction task is divided into multiple sharding tasks according to the extraction sharding field in the extraction configuration parameters, including: dividing the data extraction task into multiple sharding tasks according to the minimum value, maximum value and preset number of tasks of the extraction sharding field in the extraction configuration parameters; constructing a task directed acyclic graph according to the multiple sharding tasks, and the task directed acyclic graph is used to indicate the execution order of the sharding tasks.

[0071] For example, the system will split the data extraction task according to the minimum and maximum values ​​of the extracted shard fields and the preset number of tasks. For example, first determine the value range of the shard field (i.e., the minimum and maximum values), and then evenly split the data into multiple shards according to the preset number of tasks. For example, the splitting strategy can be dynamically adjusted according to factors such as the actual data volume and system load to ensure the execution efficiency and load balance of the sharding task.

[0072] For example, after splitting multiple shard tasks, the system will build a directed acyclic graph (DAG) of tasks based on the dependencies between these tasks. DAG is a directed graph in which each node represents a task and directed edges represent the dependencies between tasks. DAG not only represents the dependencies between tasks, but also indicates the execution order of tasks. A task can only be executed when all its predecessor tasks have been completed. This ensures the correctness and order of data extraction tasks.

[0073] In some embodiments, the data extraction task is divided into a plurality of sharding tasks according to the minimum value, the maximum value and the preset number of tasks of the extraction sharding field in the extraction configuration parameter, including:

[0074] Generate structured query language statements corresponding to the data extraction tasks for different source database types, and query the minimum value, maximum value and preset number of tasks of the extracted shard fields;

[0075] If the maximum value and the minimum value are not null values, the structured query language statement is segmented according to the minimum value, the maximum value and the preset number of tasks of the extracted shard field in the extraction configuration parameters to construct a plurality of shard intervals, each shard interval corresponding to a data range at one end;

[0076] Based on the substructured query language statement generated for each shard interval, each shard task is determined, and the multiple shard intervals are traversed to obtain multiple shard tasks.

[0077] In some embodiments, the method further includes: if at least one of the maximum value and the minimum value is a null value, the structured query language statement corresponding to the data extraction task is not split, so as to determine the data extraction task as a single-shard task with a task number of 1.

[0078] For example, first, the structured query language (SQL) statements for the corresponding data extraction tasks are generated according to different source database types (such as Hive, MySQL, Oracle, etc.). These SQL statements are used to extract data from the source database. The generated SQL statements need to include subqueries for querying the minimum and maximum values ​​of the extracted shard fields (usually primary keys or unique fields). This step is to determine the scope of data extraction.

[0079] After obtaining the minimum and maximum values ​​of the extracted shard fields, you need to check whether these two values ​​are empty (NULL). If at least one of the maximum or minimum values ​​is empty, it means that the data range is incomplete or the data distribution does not meet expectations, and it is not suitable for sharding at this time. If the maximum or minimum value contains NULL, it is decided not to split the original data extraction SQL statement, but to treat the entire data extraction task as a single shard task, and the number of tasks is set to 1.

[0080] If both the maximum and minimum values ​​are not empty, the start and end values ​​of each sharding interval are calculated based on these two boundary values ​​and the preset number of tasks (that is, how many sharding tasks you want to divide the data into). These sharding intervals will cover the entire data range, and each interval corresponds to a continuous piece of data. For each sharding interval, a corresponding sub-SQL statement is generated based on the range condition of the sharding field. These sub-SQL statements will be used to perform the actual data extraction operation, and each sub-SQL statement corresponds to a sharding task.

[0081] By traversing all the calculated sharding intervals, corresponding sub-SQL statements are generated for each interval, and the specific content of each sharding task is determined. Then, according to the number of sharding intervals and the sub-SQL statements of each interval, multiple sharding tasks are obtained. These tasks can be executed in parallel to improve the efficiency of data extraction.

[0082] Through the above steps, the data extraction task can be intelligently divided into multiple shard tasks according to the minimum and maximum values ​​of the extraction shard field in the extraction configuration parameters and the preset number of tasks. When the data range is incomplete or does not meet expectations, the system can automatically adjust the strategy and treat the task as a single shard task. This flexible data extraction strategy helps improve the efficiency and reliability of data processing.

[0083] In some embodiments, constructing a directed acyclic graph of tasks according to the plurality of sharding tasks includes:

[0084] Based on the number of the multiple shard tasks and the preset number of parallel threads, construct the multiple shard tasks into a task directed acyclic graph;

[0085] Wherein, if the data extraction task has a historical execution record, then based on the historical execution time of each shard task, predict the task arrangement with the shortest overall time consumption, and construct the task directed acyclic graph based on the task arrangement with the shortest overall time consumption; or

[0086] If the data extraction task has no historical execution record, a directed acyclic graph of the task is constructed based on the random arrangement of the sharding tasks.

[0087] For example, the system will decide how to organize these shard tasks into an ordered execution graph based on the number of multiple shard tasks and the preset number of parallel threads (that is, the upper limit of the number of tasks that can be executed simultaneously).

[0088] like Figure 3 As shown, for example, the number of parallel threads is 4, including thread 1 (Thread1), thread 2 (Thread2), thread 3 (Thread3) and thread 4 (Thread4), and each thread can execute 3 slice tasks. For example, thread 1 (Thread1) is used to execute slice task 1 (Task1), slice task 5 (Task5) and slice task 9 (Task9), and thread 2 (Thread2) is used to execute slice task 2 (Task2), slice task 6 (Task6) and slice task 10 (Task10).

[0089] For example, each shard task is considered a node in a directed acyclic (DAG) graph, and the dependencies between nodes (i.e., which tasks need to be executed first and which tasks can be executed in parallel) are determined based on the logic of data extraction and the need for parallel execution. Then, these nodes and dependencies together form a DAG graph.

[0090] Among them, if the data extraction task has historical execution records, the system can predict the overall time consumption under different task arrangements based on the historical execution time of each shard task in these records. Through algorithmic analysis, a task arrangement method can be found so that the total execution time of all shard tasks is the shortest under this method, that is, the task arrangement method with the shortest overall time consumption is predicted. This method takes into account the dependencies between tasks and the possibility of parallel execution, as well as the execution time of each task. Then, based on the predicted task arrangement method with the shortest overall time consumption, a task directed acyclic graph (DAG) is constructed.

[0091] Among them, if the data extraction task has no historical execution records, the system cannot predict the pros and cons of the task arrangement based on historical data. In this case, the system can adopt a simple strategy, that is, randomly arrange the sharding tasks and build the corresponding DAG graph. Even in the absence of historical execution records, the system can still dynamically adjust the execution plan during the task execution process. For example, when a task fails or times out, the system can reallocate tasks or adjust the dependencies between tasks to ensure the smooth completion of the data extraction task.

[0092] In the task DAG diagram, the preparation work such as the establishment of the target table and the temporary table is usually regarded as the preparation phase (Prepare) task and used as the starting node. After all sharding tasks are completed, the data in the temporary table needs to be submitted to the target table. This process is usually regarded as the submission phase (Submit) task and used as the last node in the task DAG diagram.

[0093] By constructing a directed acyclic (DAG) graph of tasks, the system can execute multiple shard tasks in an orderly and efficient manner. In the case of historical execution records, the system can use these data to optimize the task arrangement, thereby shortening the overall execution time. In the absence of historical execution records, the system can still ensure the smooth completion of tasks by randomly arranging or dynamically adjusting the execution plan. This flexible task scheduling strategy helps to improve the efficiency and reliability of data extraction.

[0094] Step 140, execute each sharding task, and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches.

[0095] Before the sharding task is executed, the system will obtain the sharding task (Shard) and the corresponding replicas (Replicas) of ClickHouse. Each sharding task runs independently and processes the data range it is responsible for. In order to optimize performance and resource utilization, these tasks will batch the generated sharding extraction result data and write it to multiple temporary files.

[0096] For example, each shard task sends its corresponding sub-SQL query through the previously configured data access interface.

[0097] The data access interface extracts data from the data source (such as a relational database, NoSQL database, etc.) according to the query request, and calls back the shard extraction result data corresponding to each shard task to the shard task in batches.

[0098] In some embodiments, before writing the shard extraction result data corresponding to each shard task into multiple temporary files in batches, the method further includes:

[0099] The shard extraction result data is preprocessed to obtain preprocessed shard extraction result data to meet the storage requirements of the distributed column database management system, and the preprocessing includes at least one of data filtering and data formatting.

[0100] For example, before writing data to a temporary file, the system performs preprocessing operations. Preprocessing operations can include data filtering (such as removing duplicate data, filtering data that meets specific conditions, etc.) and data formatting (such as converting data into a format suitable for ClickHouse storage). The preprocessing stage can also include processing special characters, converting data types, etc. to ensure that the data can be successfully received and processed by ClickHouse.

[0101] After preprocessing, the data will be written into multiple temporary files by the sharding task in batches. For example, the temporary file can be a CSV file in character-separated value (Comma-Separated Values, CSV) format. These CSV files are used as temporary storage for subsequent data upload processes.

[0102] The number of CSV files matches the number of replica nodes in the shard task. Each CSV file corresponds to a replica node, which ensures that the data can be evenly distributed to each node when uploaded to ClickHouse, thereby reducing the pressure on a single node and improving overall performance.

[0103] The writing thread is responsible for organizing the data into CSV format and continuously writing it to the corresponding file. This process is performed concurrently to further improve the efficiency of data processing.

[0104] The system records the path, size, data range, and other information of each CSV file so that the data can be correctly uploaded to ClickHouse in subsequent steps. At the same time, the system also monitors the storage of temporary files to ensure that data is not lost or processing is interrupted due to problems such as insufficient disk space.

[0105] Through this step, the system can efficiently organize the shard extraction result data into a format suitable for ClickHouse storage and prepare for the subsequent data upload process. This not only improves the efficiency of data processing, but also ensures the integrity and accuracy of the data.

[0106] In some embodiments, executing each shard task includes:

[0107] Each shard task is executed according to the task directed acyclic graph.

[0108] For example, before executing each shard task, the system will parse the constructed task directed acyclic graph to determine the execution order and dependencies of each task. During the parsing process, the system will check for circular dependencies between tasks. If there is a circular dependency, an error will be reported and the execution of the task will be terminated.

[0109] Then, according to the parsed task directed acyclic graph, the system will schedule tasks in the predetermined order and execute each shard task. During the execution process, the system will ensure that the predecessor tasks of each shard task have been successfully completed to avoid data inconsistency or task failure.

[0110] For each sharding task, the system calls the corresponding data access interface, sends the sub-SQL query, and processes the callback data. Then, the data is written into multiple temporary files in a predetermined format to prepare for the subsequent upload process.

[0111] For example, during the execution of sharding tasks, the system will monitor the execution status, progress, and results of each sharding task in real time. If a sharding task fails or an abnormal situation occurs, the system will promptly record the error information and take appropriate measures to recover or retry. At the same time, the system will also record the execution log of the sharding task for subsequent analysis and troubleshooting.

[0112] Step 150 : compress the target temporary files whose file sizes reach a preset threshold among the multiple temporary files and asynchronously upload them to the temporary table.

[0113] Step 150 ensures the efficiency and data integrity of the data extraction process by asynchronously compressing and uploading temporary files, and properly managing data writing and file cleaning.

[0114] In some embodiments, compressing the target temporary file whose file size among the multiple temporary files reaches a preset threshold and asynchronously uploading it to the temporary table includes:

[0115] Compressing a target temporary file whose file size among the multiple temporary files reaches a preset threshold through an asynchronous thread to obtain a compressed target temporary file, wherein the target temporary file corresponds to a target replica node in the distributed column database management system;

[0116] The compressed target temporary file is asynchronously uploaded to the second local table of the target replica node through the Hypertext Transfer Protocol port of the target replica node.

[0117] For example, the system detects the sizes of multiple temporary files, and once the size of a temporary file (such as a CSV file) reaches a preset threshold, it is marked as a target temporary file.

[0118] Then, the system starts an asynchronous thread to compress the target temporary file. The compression algorithm can be selected according to the actual situation. For example, the LZ4 algorithm is well-known for its high compression rate and fast decompression ability, and is very suitable for compressing large data files. After the compression is completed, the compressed target temporary file is obtained. This file will be used for subsequent asynchronous upload operations.

[0119] The compressed target temporary file needs to be uploaded to the target replica node in a distributed columnar database management system (such as ClickHouse). Then, the compressed target temporary file is asynchronously uploaded to the second local table of the node through the Hypertext Transfer Protocol (HTTP) port of the target replica node. The asynchronous upload operation means that the system can continue to process other tasks while uploading files, such as receiving new data, processing other temporary files, etc., thereby improving the system's concurrent processing capabilities and response speed.

[0120] During the asynchronous upload of CSV files, in order to ensure the continuity and integrity of the data, the system will write the subsequent data into a new CSV file. For example, when CSV-A-part.0 is being uploaded, the subsequent data will be written into CSV-A-part.1.

[0121] For example, once a CSV file (the target temporary file) is uploaded, the system will immediately clean up the file to free up disk space and avoid disk accumulation. This helps keep the system clean and running efficiently.

[0122] When the data reading of the sharding task is completed, the system will check and upload all remaining CSV files that have data written but not uploaded.

[0123] The sharding task is officially completed only when all CSV files are successfully uploaded. This means that the sharding extraction result data of the sharding task has been safely stored in the distributed column database management system and can be processed and analyzed at any time.

[0124] like Figure 4 The execution diagram of shard task 1 (Shard 1) shown in the figure reads the shard extraction result data corresponding to shard task 1 (Shard 1) through a data reading operation; then performs a data preprocessing operation (such as data filtering) on ​​the shard extraction result data to obtain the preprocessed shard extraction result data, and then performs data splitting to write the split data into multiple temporary files (such as CSV-A-part.1, CSV-B-part.1, CSV-C-part.1, CSV-D-part.1) in batches. For example, the target replica node corresponding to the temporary file CSV-A-part.1 is replica node 1 (Replicas 1), the target replica node corresponding to the temporary file CSV-B-part.1 is replica node 2 (Replicas 2), the target replica node corresponding to the temporary file CSV-C-part.1 is replica node 3 (Replicas 3), and the target replica node corresponding to the temporary file CSV-D-part.1 is replica node 4 (Replicas 4).

[0125] For example, the compressed temporary file CSV-A-part.1 is asynchronously uploaded to the second local table (such as ClickHouse Local Table) through the HTTP port of replica node 1 (Replicas 1).

[0126] For example, the compressed temporary file CSV-B-part.1 is asynchronously uploaded to the second local table (such as ClickHouse Local Table) through the HTTP port of replica node 2 (Replicas 2).

[0127] For example, the compressed temporary file CSV-C-part.1 is asynchronously uploaded to the second local table (such as ClickHouse Local Table) through the HTTP port of replica node 3 (Replicas 3).

[0128] For example, the compressed temporary file CSV-D-part.1 is asynchronously uploaded to the second local table (such as ClickHouse Local Table) through the HTTP port of replica node 4 (Replicas 4).

[0129] Step 160, when all sharding tasks are completed, the data in the temporary table is submitted to the target table to complete the data extraction.

[0130] When all sharding tasks are completed, the system triggers a data submission operation to submit the data in the temporary table to the target table. This means that all sharding tasks have successfully written data to each local table instance of the temporary table.

[0131] For example, for a temporary table shadow_A of a data extraction task, it will correspond to a second local table shadow_A_local on each replica (ClickHouse) node; similarly, the target table target_A will also correspond to a first local table target_A_local on each node.

[0132] In ClickHouse, the EXCHANGE TABLES statement is used to exchange the metadata of two tables, but does not affect the data in the tables. This means that the table names and the locations of the data files will be swapped, but the data content remains unchanged. In an embodiment of the present application, the role of the EXCHANGE TABLES statement is to exchange the table names of the second local table (such as shadow_A_local) and the first local table (such as target_A_local) on each node in the ClickHouse Cluster. In this way, the data originally stored in the temporary table will be "migrated" to the target table.

[0133] In some embodiments, the method further comprises:

[0134] The temporary table and the second local table of all replica nodes are deleted.

[0135] After the EXCHANGE TABLES statement is successfully executed, the system deletes the original temporary table (shadow_A) and its second local table instance (shadow_A_local) on each node to ensure data uniqueness and avoid confusion.

[0136] After the above operations, the target table (target_A) and its first local table instance (target_A_local) on each node will be retained and contain the complete data of this data extraction.

[0137] For example, Figure 4As shown in the figure, the processing progress corresponding to the data extraction task can also be pushed to the business platform. During the execution of the data extraction task, the system will continuously record the processing progress of each stage, such as the completion of parameter configuration, task segmentation, task scheduling, sharding task execution, data submission and other steps. These progress information may include key indicators such as the amount of data processed, processing speed, and remaining time estimation. The system will encapsulate the captured progress information into a structured data format (such as JSON, XML, etc.) for transmission through API, Web service and other interfaces. The encapsulated progress information will be sent to the business platform in real time or periodically through a preset push mechanism (such as HTTP POST request, WebSocket connection, etc.). The business platform will receive the progress information pushed by the system and display it on the graphical user interface, such as a progress bar, percentage indicator, detailed report, etc. The business party can intuitively understand the current status and overall progress of the data extraction task through the graphical user interface.

[0138] All of the above technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described in detail here.

[0139] The embodiment of the present application receives a data extraction task and obtains the extraction configuration parameters corresponding to the data extraction task; according to the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in a distributed column database management system; according to the extraction configuration parameters, the data extraction task is divided into multiple sharding tasks; each sharding task is executed, and the sharding extraction result data corresponding to each sharding task is written in batches to multiple temporary files; the target temporary files whose file sizes reach a preset threshold in multiple temporary files are compressed and asynchronously uploaded to the temporary table; when all sharding tasks are executed, the data in the temporary table is submitted to the target table to complete the data extraction. The embodiment of the present application realizes efficient, accurate and flexible data extraction operations by adopting advanced technical means such as distributed processing, parallel execution of sharding tasks, and asynchronous uploading, which significantly improves data processing efficiency and system resource utilization. Among them, the extraction performance is improved by multi-level splitting of the sharding extraction result data and asynchronous batch uploading, and the data storage structure is optimized by fine extraction configuration parameters, thereby improving query efficiency.

[0140] In order to better implement the data extraction method of the embodiment of the present application, the embodiment of the present application also provides a data extraction device. Figure 5 , Figure 5 The data extraction device 200 is a schematic diagram of the structure of the data extraction device provided in the embodiment of the present application. The data extraction device 200 may include:

[0141] The acquisition unit 210 is used to receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task;

[0142] An establishing unit 220, configured to establish a target table and a temporary table corresponding to the data extraction task in a distributed column-based database management system according to the extraction configuration parameters;

[0143] A splitting unit 230, configured to split the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters;

[0144] An execution unit 240 is used to execute each sharding task and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches;

[0145] The processing unit 250 is used to compress the target temporary files whose file sizes reach a preset threshold among the multiple temporary files and asynchronously upload them to the temporary table;

[0146] The submitting unit 260 is used to submit the data in the temporary table to the target table to complete data extraction after all sharding tasks are executed.

[0147] In some embodiments, the data extraction device 200 further includes:

[0148] A configuration unit is used to configure the extraction configuration parameters, wherein the extraction configuration parameters include at least one of the extraction fragmentation field, partition field, calculation primary key and field non-null value identification information.

[0149] In some embodiments, the segmentation unit 230 may be used to segment the data extraction task into a plurality of segmentation tasks according to the extraction segmentation field in the extraction configuration parameters.

[0150] In some embodiments, when the segmentation unit 230 segments the data extraction task into a plurality of segmentation tasks according to the extraction segmentation field in the extraction configuration parameter, it can be used to:

[0151] According to the minimum value, the maximum value and the preset number of tasks of the extraction shard field in the extraction configuration parameters, the data extraction task is divided into a plurality of shard tasks;

[0152] A task directed acyclic graph is constructed according to the multiple sharding tasks, where the task directed acyclic graph is used to indicate the execution order of the sharding tasks.

[0153] In some embodiments, when the splitting unit 230 splits the data extraction task into a plurality of sharding tasks according to the minimum value, the maximum value and the preset number of tasks of the extraction sharding field in the extraction configuration parameters, it can be used to:

[0154] Generate structured query language statements corresponding to the data extraction tasks for different source database types, and query the minimum value, maximum value and preset number of tasks of the extracted shard fields;

[0155] If the maximum value and the minimum value are not null values, the structured query language statement is segmented according to the minimum value, the maximum value and the preset number of tasks of the extracted shard field in the extraction configuration parameters to construct a plurality of shard intervals, each shard interval corresponding to a data range at one end;

[0156] Based on the substructured query language statement generated for each shard interval, each shard task is determined, and the multiple shard intervals are traversed to obtain multiple shard tasks.

[0157] In some embodiments, the splitting unit 230 can also be used for: if at least one of the maximum value and the minimum value is a null value, the structured query language statement corresponding to the data extraction task is not split, so as to determine the data extraction task as a single-shard task with a task number of 1.

[0158] In some embodiments, when constructing a task directed acyclic graph according to the plurality of sharding tasks, the slicing unit 230 may be used to:

[0159] Based on the number of the multiple shard tasks and the preset number of parallel threads, construct the multiple shard tasks into a task directed acyclic graph;

[0160] Wherein, if the data extraction task has a historical execution record, then based on the historical execution time of each shard task, predict the task arrangement with the shortest overall time consumption, and construct the task directed acyclic graph based on the task arrangement with the shortest overall time consumption; or

[0161] If the data extraction task has no historical execution record, a directed acyclic graph of the task is constructed based on the random arrangement of the sharding tasks.

[0162] In some embodiments, the execution unit 240 may be used to:

[0163] Each shard task is executed according to the task directed acyclic graph.

[0164] In some embodiments, the establishing unit 220 may be used to:

[0165] According to the partition field in the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in the distributed column-based database management system, wherein the target table is configured with a first local table of each replica node in the distributed column-based database management system, and the temporary table is configured with a second local table of each replica node in the distributed column-based database management system, and each temporary file corresponds to a replica node in the distributed column-based database management system.

[0166] In some embodiments, the processing unit 250 may be configured to:

[0167] Compressing a target temporary file whose file size among the multiple temporary files reaches a preset threshold through an asynchronous thread to obtain a compressed target temporary file, wherein the target temporary file corresponds to a target replica node in the distributed column database management system;

[0168] The compressed target temporary file is asynchronously uploaded to the second local table of the target replica node through the Hypertext Transfer Protocol port of the target replica node.

[0169] In some embodiments, the data extraction device 200 further includes:

[0170] The deleting unit is used to delete the temporary table and the second local tables of all replica nodes.

[0171] In some embodiments, before writing the shard extraction result data corresponding to each shard task into multiple temporary files in batches, the execution unit 240 may also be used to:

[0172] The shard extraction result data is preprocessed to obtain preprocessed shard extraction result data to meet the storage requirements of the distributed column database management system, and the preprocessing includes at least one of data filtering and data formatting.

[0173] Each unit in the above data extraction device can be implemented in whole or in part by software, hardware, or a combination thereof. Each unit can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each unit.

[0174] The data extraction device 200 may be integrated in a terminal or server that has a storage device and a processor and has computing capabilities, or the data extraction device 200 is the terminal or server.

[0175] Optionally, the present application further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0176] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a terminal or a server. Figure 6 As shown, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303 and a communication bus 304. The communication interface 301, the memory 302, and the processor 303 communicate with each other through the communication bus 304. The communication interface 301 is used for the computer device 300 to communicate data with external devices. The memory 302 can be used to store software programs and modules, and the processor 303 runs the software programs and modules stored in the memory 302, such as the software programs of the corresponding operations in the aforementioned method embodiments.

[0177] Optionally, the processor 303 may call the software program and module stored in the memory 302 to perform the following operations:

[0178] Receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task; establish a target table and a temporary table corresponding to the data extraction task in a distributed column-based database management system according to the extraction configuration parameters; divide the data extraction task into multiple sharding tasks according to the extraction configuration parameters; execute each sharding task, and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches; compress the target temporary files whose file size reaches a preset threshold among the multiple temporary files and upload them asynchronously to the temporary table; when all sharding tasks are executed, submit the data in the temporary table to the target table to complete the data extraction.

[0179] The present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to a computer device, and the computer program enables the computer device to execute the corresponding process in the data extraction method in the embodiment of the present application, which will not be described in detail for the sake of brevity.

[0180] A schematic diagram of a storage medium provided in an embodiment of the present application, such as Figure 7As shown, a program product 400 for implementing the above method according to an exemplary embodiment of the present application is described, which can adopt a portable compact disk read-only memory (CDROM) and include program code, and can be run on a computer device, such as a mobile phone. However, the program product of the present application is not limited thereto. In the present application, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0181] The present application also provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the data extraction method in the embodiment of the present application, which will not be described here for brevity.

[0182] The present application also provides a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding process in the data extraction method in the present application, which will not be described here for the sake of brevity.

[0183] It should be understood that the processor of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and performed. The software module may be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0184] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0185] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0186] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0187] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0188] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0189] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] In addition, each functional unit in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0191] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer or a server) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0192] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A data extraction method, characterized in that: The method comprises: Receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task; According to the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in a distributed column-based database management system; According to the extraction configuration parameters, the data extraction task is divided into a plurality of sharding tasks; Execute each sharding task and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches; Compressing a target temporary file whose file size reaches a preset threshold among the multiple temporary files and asynchronously uploading it to the temporary table; When all sharding tasks are completed, the data in the temporary table is submitted to the target table to complete the data extraction.

2. The data extraction method according to claim 1, characterized in that: Before receiving the data extraction task, the method further includes: The extraction configuration parameters are configured, wherein the extraction configuration parameters include at least one of extraction of fragmentation fields, partition fields, calculation of primary keys, and field non-null value identification information.

3. The data extraction method according to claim 2, characterized in that: The step of dividing the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters includes: According to the extraction sharding field in the extraction configuration parameters, the data extraction task is divided into a plurality of sharding tasks.

4. The data extraction method according to claim 3, characterized in that: The step of dividing the data extraction task into a plurality of sharding tasks according to the extraction sharding field in the extraction configuration parameter comprises: According to the minimum value, the maximum value and the preset number of tasks of the extraction shard field in the extraction configuration parameters, the data extraction task is divided into a plurality of shard tasks; A task directed acyclic graph is constructed according to the multiple sharding tasks, where the task directed acyclic graph is used to indicate the execution order of the sharding tasks.

5. The data extraction method according to claim 4, characterized in that: The step of dividing the data extraction task into a plurality of sharding tasks according to the minimum value, the maximum value and the preset number of tasks of the extraction sharding field in the extraction configuration parameter comprises: Generate structured query language statements corresponding to the data extraction tasks for different source database types, and query the minimum value, maximum value and preset number of tasks of the extracted shard fields; If the maximum value and the minimum value are not null values, the structured query language statement is segmented according to the minimum value, the maximum value and the preset number of tasks of the extracted shard field in the extraction configuration parameters to construct a plurality of shard intervals, each shard interval corresponding to a data range at one end; Based on the substructured query language statement generated for each shard interval, each shard task is determined, and the multiple shard intervals are traversed to obtain multiple shard tasks.

6. The data extraction method according to claim 2, characterized in that: The step of establishing a target table and a temporary table corresponding to the data extraction task in a distributed column-based database management system according to the extraction configuration parameters includes: According to the partition field in the extraction configuration parameters, a target table and a temporary table corresponding to the data extraction task are established in the distributed column-based database management system, wherein the target table is configured with a first local table of each replica node in the distributed column-based database management system, and the temporary table is configured with a second local table of each replica node in the distributed column-based database management system, and each temporary file corresponds to a replica node in the distributed column-based database management system.

7. A data extraction device, characterized in that: The device comprises: An acquisition unit, used to receive a data extraction task and obtain extraction configuration parameters corresponding to the data extraction task; An establishing unit, used to establish a target table and a temporary table corresponding to the data extraction task in a distributed column database management system according to the extraction configuration parameters; A splitting unit, used for splitting the data extraction task into a plurality of sharding tasks according to the extraction configuration parameters; An execution unit, used to execute each sharding task and write the sharding extraction result data corresponding to each sharding task into multiple temporary files in batches; A processing unit, configured to compress a target temporary file whose file size reaches a preset threshold among the multiple temporary files and asynchronously upload the target temporary file to the temporary table; The submitting unit is used to submit the data in the temporary table to the target table to complete data extraction after all sharding tasks are executed.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the data extraction method according to any one of claims 1 to 6.

9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the data extraction method according to any one of claims 1 to 6 by calling the computer program stored in the memory.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the data extraction method according to any one of claims 1 to 6 is implemented.