Breakpoint resuming method, apparatus, device, and medium
Patent Information
- Application Number
- CN202311661731.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-05
AI Technical Summary
[0003]目前再将数据文件导入数据库时,每一批次提交后会写一个输出文件,输出文件中存储了作业起始点至作业中断点的作业数据,但是基于数据文件实现断点重提,文件读写涉及输入/输出(Input/Output,简称I/O)接口,但是I/O接口通常耗时较长,并且将数据文件拆分成多个小文件并基于多线程同时处理时,会造成对输出文件的抢占情况和数据一致性问题,造成断点重提与将数据文件导入数据库的效率被严重降低
[0035]本申请提供的断点重提方法、装置、设备及介质,对原始数据文件进行拆分得到多个中间数据文件,并将每个中间数据文件按照预设的数据量大小进行拆分得到对应的多个待处理文件;通过多个线程异步处理的方式,将每个待处理文件插入到目标数据库中,每个线程处理一个待处理文件;将所述原始数据文件的拆分信息记录在Redis数据库中,所述原始数据文件的拆分信息包括原始数据文件的拆分断点和已拆分得到的待处理文件的文件信息;在需要断点重提时,从所述Redis数据库中获取所述原始数据文件的拆分信息,并基于所述拆分信息继续完成对所述原始数据文件的后续处理。基于本申请提供的方法,通过多线程异步处理所述多个待处理文件,并通过Redis数据库记录所述原始数据文件的拆分信息,不存在对所述Redis数据库的抢占情况,可保障拆分信息的准确记录,并在需要提取数据文件的拆分信息继续处理所述数据文件时,可及时的从所述Redis数据库中提取拆分信息,极大的提高了断点重提并将数据文件导入数据库的效率。
Smart Images

Figure CN117633084B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular to a method, apparatus, device, and medium for retrieving breakpoints. Background Technology
[0002] During the process of importing data files into the database, there may be situations where the job is interrupted due to abnormal reasons such as environment or resources. If the job is re-executed from the beginning the next time it is re-submitted, the resources will be wasted. By recording the job program breakpoint, the job program can be restarted from the recorded breakpoint when it is re-executed.
[0003] Currently, when importing data files into the database, an output file is written after each batch submission. The output file stores the job data from the job start point to the job breakpoint. However, implementing breakpoint retrieval based on data files involves input / output (I / O) interfaces, which are usually time-consuming. Furthermore, splitting the data file into multiple small files and processing them simultaneously using multiple threads can cause preemption of output files and data consistency issues, severely reducing the efficiency of breakpoint retrieval and importing data files into the database. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for retrieving breakpoints, which can improve the efficiency of retrieving breakpoints and importing data files into a database.
[0005] On the one hand, this application provides a breakpoint re-capture method, the method comprising:
[0006] The original data file is split into multiple intermediate data files, and each intermediate data file is further split into multiple files to be processed according to a preset data size.
[0007] By using multiple threads to process asynchronously, each file to be processed is inserted into the target database, with each thread processing one file to be processed.
[0008] The splitting information of the original data file is recorded in the Redis database. The splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split files to be processed.
[0009] The splitting information of the original data file is obtained from the Redis database, and the subsequent processing of the original data file is carried out based on the splitting information.
[0010] In one example, splitting the original data file into multiple intermediate data files includes:
[0011] The original data file is split into data fragments according to the splitting rules of the fragment key, and the fragment key corresponds to a data node;
[0012] For each data slice obtained, add a corresponding current file write pointer, and write the data slice to the intermediate data file corresponding to the current file write pointer according to the file write pointer;
[0013] When the amount of data in the data fragments written to the intermediate data file reaches a preset first threshold, the current file write pointer is updated.
[0014] In one example, recording the splitting information of the original data file in the Redis database includes:
[0015] After each file to be processed is split, the file information of the split file to be processed is recorded in the Redis database;
[0016] Once the thread starts importing the data of the file to be processed into the target database, it sends a feedback to the Redis database so that the Redis database updates the file information of the file to be processed to indicate that the process is complete.
[0017] In one example, obtaining the splitting information of the original data file from the Redis database and continuing the subsequent processing of the original data file based on the splitting information includes:
[0018] Determine if a split breakpoint for the original data file exists in the Redis database. If it exists, continue splitting the original data file based on the split breakpoint. If it does not exist, determine if file information for the file to be processed exists in the Redis database. If it exists, continue processing the file to be processed using multi-threading based on the file information. If it does not exist, directly split the original data file.
[0019] In one example, inserting each file to be processed into the target database using a multi-threaded asynchronous processing method includes:
[0020] It is determined whether the number of files to be processed has reached a preset second threshold. If the number of files to be processed has reached the second threshold, each file to be processed is inserted into the target database asynchronously through multiple threads using JDBC batch processing.
[0021] In one example, the method includes:
[0022] Once all the file information for all pending files related to the original data file in the splitting information recorded in the Redis database has been completed, clear the splitting information for the original data file recorded in the Redis database.
[0023] On the other hand, this application provides a breakpoint recovery device, the device comprising:
[0024] The splitting module is used to split the original data file into multiple intermediate data files, and then split each intermediate data file into multiple files to be processed according to a preset data size.
[0025] The insert module is used to insert each file to be processed into the target database in an asynchronous manner using multiple threads, with each thread processing one file to be processed.
[0026] The recording module is used to record the splitting information of the original data file in the Redis database. The splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split files to be processed.
[0027] The processing module is used to obtain the splitting information of the original data file from the Redis database, and to continue to complete the subsequent processing of the original data file based on the splitting information.
[0028] In one example, the splitting module is specifically used to split the original data file according to the splitting rules of the splitting key to obtain data fragments, wherein the splitting key corresponds to a data node;
[0029] The splitting module is further configured to add a corresponding current file write pointer to each obtained data slice, and write the data slice to the intermediate data file corresponding to the current file write pointer according to the file write pointer;
[0030] The splitting module is further configured to update the current file write pointer when the amount of data in the data shards written to the intermediate data file reaches a preset first threshold.
[0031] In another aspect, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0032] The memory stores computer-executed instructions;
[0033] The processor executes computer execution instructions stored in the memory to implement the method described above.
[0034] In another aspect, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in the preceding claim.
[0035] The breakpoint retrieval method, apparatus, device, and medium provided in this application split the original data file into multiple intermediate data files, and further split each intermediate data file into multiple files to be processed according to a preset data size. Each file to be processed is inserted into a target database through asynchronous processing by multiple threads, with each thread processing one file. The splitting information of the original data file is recorded in a Redis database, including the splitting breakpoint of the original data file and the file information of the split files to be processed. When breakpoint retrieval is needed, the splitting information of the original data file is retrieved from the Redis database, and subsequent processing of the original data file is continued based on the splitting information. Based on the method provided in this application, the multiple files to be processed are processed asynchronously by multiple threads, and the splitting information of the original data file is recorded in the Redis database. There is no contention for the Redis database, ensuring accurate recording of the splitting information. When the splitting information of a data file needs to be retrieved to continue processing, the splitting information can be retrieved from the Redis database in a timely manner, greatly improving the efficiency of breakpoint retrieval and importing the data file into the database. Attached Figure Description
[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0037] Figure 1 This is a schematic diagram illustrating an application scenario for this application.
[0038] Figure 2 This is a flowchart illustrating a breakpoint re-retrieval method provided in Embodiment 1 of this application;
[0039] Figure 3 A flowchart illustrating another breakpoint re-retrieval method provided in Embodiment 1 of this application;
[0040] Figure 4 A flowchart illustrating another breakpoint re-retrieval method provided in Embodiment 1 of this application;
[0041] Figure 5 This is a flowchart illustrating the process of breakpoint re-retrieval and breakpoint update as an example.
[0042] Figure 6This is a schematic diagram of a breakpoint re-lifting device provided in Embodiment 2 of this application;
[0043] Figure 7 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application.
[0044] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0045] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0046] During the process of importing data files into the database, there may be situations where the job is interrupted due to abnormal reasons such as environment or resources. If the job is re-executed from the beginning the next time it is re-submitted, the resources will be wasted. By recording the job program breakpoint, the job program can be restarted from the recorded breakpoint when it is re-executed.
[0047] Currently, when importing data files into the database, an output file is written after each batch submission. The output file stores the job data from the job start point to the job breakpoint. However, implementing breakpoint retrieval based on data files involves input / output (I / O) interfaces, which are usually time-consuming. Furthermore, splitting the data file into multiple small files and processing them simultaneously using multiple threads can cause preemption of output files and data consistency issues, severely reducing the efficiency of breakpoint retrieval and importing data files into the database.
[0048] Figure 1This is an example application scenario illustration of this application. The original data file is split into multiple intermediate data files, and each intermediate data file is further split into multiple files to be processed according to a preset data size. Each file to be processed is inserted into the target database using a multi-threaded asynchronous processing method, with each thread processing one file. The splitting information of the original data files is recorded in a Redis database. When a breakpoint retrieval is required, the splitting information of the original data files is retrieved from the Redis database, and subsequent processing of the original data files continues based on this information. Based on the method provided in this application, the multiple files to be processed are processed asynchronously using multiple threads, and the splitting information of the original data files is recorded in a Redis database. There is no contention for the Redis database, ensuring accurate recording of splitting information. Furthermore, when the splitting information of a data file needs to be retrieved to continue processing, it can be retrieved from the Redis database in a timely manner, greatly improving the efficiency of breakpoint retrieval and importing the data file into the database.
[0049] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0050] The technical solutions of this application will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In the description of this application, unless otherwise expressly specified and limited, the terms should be broadly understood within the art. The embodiments of this application will now be described with reference to the accompanying drawings.
[0051] Example 1
[0052] Figure 2 This is a flowchart illustrating a breakpoint re-capture method provided in Embodiment 1 of this application, as shown below. Figure 2 As shown, the method includes:
[0053] Step 201: Split the original data file to obtain multiple intermediate data files, and split each intermediate data file according to the preset data size to obtain multiple corresponding files to be processed;
[0054] Step 202: Insert each file to be processed into the target database using a multi-threaded asynchronous processing method, with each thread processing one file to be processed.
[0055] Step 203: Record the splitting information of the original data file in the Redis database. The splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split files to be processed.
[0056] Step 204: Obtain the splitting information of the original data file from the Redis database, and continue to complete the subsequent processing of the original data file based on the splitting information.
[0057] The execution entity in this embodiment is a breakpoint re-emergence device, which can be implemented by a computer program, such as application software; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive or cloud drive; or it can be implemented by a physical device that integrates or installs relevant computer programs, such as a chip.
[0058] In conjunction with the scenario example, splitting the original data file is mainly to improve the processing efficiency of the original data file. The splitting of the original data file is mainly divided into two steps. First, the original data file is split into multiple intermediate data files. The size of the intermediate data files is limited to avoid the intermediate data files being too large. Second, each intermediate data file is split into multiple files to be processed according to a preset data volume. The thread pool contains multiple threads, each responsible for processing one file to be processed. When the number of files to be processed reaches a certain amount, the multiple threads in the thread pool batch insert each file into the target database at once. Specifically, the files to be processed are added to a task queue, and then the multiple threads in the thread pool sequentially retrieve one file from the task queue and batch insert it into the target database using Java Database Connectivity (JDBC). The target database can be a Java database. The JDBC batch processing refers to combining related Structured Query Language (SQL) statements into a batch and submitting it within a Java database connection.
[0059] The Redis database records file information of the file to be processed and the splitting breakpoints of the original data file. These splitting breakpoints primarily record the splitting progress of the original data file. On one hand, the splitting breakpoints of the original data file can be used as the splitting information of the original data file; on the other hand, the recorded file information of the file to be processed can be used as the splitting information of the original data file. The splitting breakpoints are generally recorded line by line. The file information of the file to be processed includes two categories: completed and incomplete. If the file information of the file to be processed is incomplete, it means that the file to be processed has not been successfully inserted into the target database. If the file information of the file to be processed is completed, it means that the file to be processed has been successfully inserted into the target database. The splitting information of the original data file is retrieved from the Redis database. If the splitting information is a splitting breakpoint of the original data file, the splitting of the original data file continues based on the splitting breakpoint. If the splitting information is an incomplete file information of the file to be processed, the file to be processed is re-inserted into the target database via a thread.
[0060] Based on the method provided in this example, the multiple files to be processed are processed asynchronously through multi-threading, and the splitting information of the original data files is recorded in the Redis database. There is no preemption of the Redis database, which can ensure the accurate recording of splitting information. When it is necessary to extract the splitting information of the data file to continue processing the data file, the splitting information can be extracted from the Redis database in a timely manner, which greatly improves the efficiency of re-extracting from the breakpoint and importing the data file into the database.
[0061] Optional, Figure 3 This is a flowchart illustrating another breakpoint re-retrieval method provided in Embodiment 1 of this application, as shown below. Figure 3 As shown, in step 201, splitting the original data file into multiple intermediate data files includes:
[0062] Step 301: Split the original data file according to the splitting rules of the splitting key to obtain data fragments, where each splitting key corresponds to a data node;
[0063] Step 302: Add a corresponding current file write pointer to each obtained data fragment, and write the data fragment to the intermediate data file corresponding to the current file write pointer according to the file write pointer;
[0064] Step 303: When the amount of data in the data fragments written to the intermediate data file reaches a preset first threshold, the current file write pointer is updated.
[0065] In this example, the sharding key typically consists of one or more column fields. Designed based on table and database sharding principles, this example uses the sharding key to determine which node each row of data in the original data file belongs to, thus allocating each row to a corresponding data shard. Each data node can be considered a data shard. Because the data size varies across data nodes, when sharding the original data file according to the sharding key, a file write pointer is added to each resulting data shard. This write pointer indicates the address to be written to next. In this example, the file write pointer added to the data shard represents the intermediate data file to which the data shard should be written. The filename of the intermediate data file includes the corresponding node information. Specifically, the current file write pointer is added to each data shard obtained from the split, so that the data shard is written to the corresponding current intermediate data file. When the amount of data written reaches the rated size, the write pointer will be updated. For example, a first threshold for the amount of data written to the intermediate data file is preset. When the amount of data written to the intermediate data file reaches the first threshold, the file write pointer of the next data file is updated to the next intermediate data file so that the next data file is written to the next intermediate data file. In this example, after the amount of data written to the intermediate data file reaches the preset threshold, updating the file write pointer of the data file can prevent the intermediate data file from becoming too large.
[0066] Optional, Figure 4 This is a flowchart illustrating another breakpoint re-retrieval method provided in Embodiment 1 of this application, as shown below. Figure 4 As shown, in step 203, recording the splitting information of the original data file in the Redis database includes:
[0067] Step 401: After each file to be processed is split, record the file information of the split files in the Redis database.
[0068] Step 402: After the thread starts importing the data of the file to be processed into the target database, it sends a feedback to the Redis database so that the Redis database updates the file information of the file to be processed to indicate that it has been completed.
[0069] In the context of a scenario example, on the one hand, the splitting information can be the splitting breakpoint of the original data file. Since the splitting of the original data file is performed according to the sharding key rule, the splitting breakpoint of the original data file is updated in the Redis database every time a split of the original data file is completed.
[0070] On the other hand, the splitting information can be the file information of the files to be processed. Specifically, when a file to be processed is first split, its file information is recorded in the Redis database. At this time, the file information of the file to be processed is "incomplete," indicating that it has not yet been inserted into the target database. When multiple files to be processed are batch-inserted into the target database using the JDBC batch processing method, because the thread is processing asynchronously, after a file to be processed is successfully inserted into the target database, the Redis database is notified of the status information that the file to be processed has been successfully inserted into the target database, and the file information of the file to be processed is changed to "complete." Recording the file information of the files to be processed in the Redis database through asynchronous processing avoids preemption of the Redis database and ensures the accurate recording of splitting information.
[0071] Optionally, step 204 includes:
[0072] Determine if a split breakpoint for the original data file exists in the Redis database. If it exists, continue splitting the original data file based on the split breakpoint. If it does not exist, determine if file information for the file to be processed exists in the Redis database. If it exists, continue processing the file to be processed using multi-threading based on the file information. If it does not exist, directly split the original data file.
[0073] Combined with scenario examples, Figure 5 The following is a flowchart illustrating the process of breakpoint re-retrieval and breakpoint update, as shown in the example. Figure 5As shown, the splitting information for the original data file can first be extracted from the Redis database. If there are no splitting breakpoints for the original data file in the Redis database, it indicates that this is the first time the task of splitting the original data file and inserting it into the target database has been executed, or the splitting has been completed. Therefore, it can be further determined whether the splitting information includes the file information of the files to be processed. If the splitting information does not include the file information of the files to be processed, it indicates that the splitting of the original data file has not yet started. In this case, the splitting work of the original data file can be directly performed, and the splitting breakpoints for the original data file recorded in the Redis database are updated based on the splitting progress. Then, each split file to be processed is inserted into the target database using JDBC multi-threading, and the file information of the files to be processed recorded in the Redis database is updated. If the splitting information includes the file information of the files to be processed, the work of inserting the files marked as incomplete in the file information into the target database continues, and the file information of the files to be processed recorded in the Redis database is updated. If the Redis database records a split breakpoint for the original initial file, then the splitting of the original data file continues based on the recorded split breakpoint.
[0074] Optionally, when all the file information of the pending files related to the original data file in the splitting information recorded in the Redis database has been completed, the splitting information of the original data file recorded in the Redis database is cleared.
[0075] In the scenario example, after each file to be processed, derived from the original data file, is inserted into the target database via a thread, it indicates that all work on the original data file is complete. Therefore, it is unnecessary to record any splitting information about the original data file in the Redis database, and the splitting information recorded in the Redis database can be cleared. At this point, all files to be processed have been successfully processed, and the intermediate data files generated during the process are also cleaned up.
[0076] Optionally, inserting each file to be processed into the target database through asynchronous processing using multiple threads includes:
[0077] It is determined whether the number of files to be processed has reached a preset second threshold. If the number of files to be processed has reached the second threshold, each file to be processed is inserted into the target database asynchronously through multiple threads using JDBC batch processing.
[0078] In the scenario example, when inserting the files to be processed into the target database using the JDBC batch processing method, JDBC connection resources need to be enabled. Since multiple threads exist in the thread pool, the JDBC connection resources can be enabled only when the number of files to be processed reaches the second threshold, thus improving the overall processing efficiency of the files. After the JDBC connection resources are enabled, the JDBC batch processing method uses multiple threads to asynchronously insert the files to be processed into the target database. Optionally, the second threshold can be selected as the number of threads in the thread pool, ensuring that each file to be processed has a corresponding thread for processing.
[0079] In this embodiment, the original data file is split into multiple intermediate data files, and each intermediate data file is further split into multiple files to be processed according to a preset data size. Each file to be processed is inserted into the target database using a multi-threaded asynchronous processing method, with each thread processing one file. The splitting information of the original data file is recorded in a Redis database. When a breakpoint retrieval is required, the splitting information of the original data file is retrieved from the Redis database, and subsequent processing of the original data file continues based on this information. Based on the method provided in this embodiment, the multiple files to be processed are processed asynchronously using multiple threads, and the splitting information of the original data file is recorded in a Redis database. There is no contention for the Redis database, ensuring accurate recording of the splitting information. Furthermore, when the splitting information of a data file needs to be retrieved to continue processing, it can be retrieved from the Redis database in a timely manner, greatly improving the efficiency of breakpoint retrieval and importing the data file into the database.
[0080] Example 2
[0081] Figure 6 This is a schematic diagram of a breakpoint re-lifting device provided in Embodiment 2 of this application, as shown below. Figure 6 As shown, the device includes:
[0082] The splitting module 61 is used to split the original data file into multiple intermediate data files, and to split each intermediate data file into multiple corresponding files to be processed according to a preset data size;
[0083] Insertion module 62 is used to insert each file to be processed into the target database in an asynchronous manner using multiple threads, with each thread processing one file to be processed.
[0084] Recording module 63 is used to record the splitting information of the original data file in the Redis database. The splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split files to be processed.
[0085] Processing module 64 is used to obtain the splitting information of the original data file from the Redis database, and to continue to complete the subsequent processing of the original data file based on the splitting information.
[0086] In conjunction with the scenario example, splitting the original data file is mainly to improve the processing efficiency of the original data file. The splitting module 61 splits the original data file in two main steps. First, the original data file is split into multiple intermediate data files. The size of the intermediate data files is limited to avoid the intermediate data files being too large. Second, each intermediate data file is split into multiple files to be processed according to a preset data volume. The thread pool contains multiple threads, each responsible for processing one file to be processed. When the number of files to be processed reaches a certain amount, the insertion module 61 uses multiple threads to batch insert each file into the target database at once. Specifically, the files to be processed are added to the task queue, and then multiple threads in the thread pool sequentially extract one file from the task queue and batch insert the files into the target database using JDBC batch processing. The target database can be a Java database. JDBC batch processing refers to combining related Structured Query Language (SQL) statements into a batch and submitting it within a Java database connection.
[0087] The recording module 63 records the splitting information of the original data file in the Redis database. This splitting information mainly records the splitting breakpoints of the original data file and the file information of the split files to be processed. On one hand, the splitting breakpoints of the original data file can be used as the splitting information of the original data file; on the other hand, the recorded file information of the files to be processed can be used as the splitting information of the original data file. The splitting breakpoints are generally recorded line by line. The file information of the files to be processed includes two categories: completed and incomplete. If the file information of the files to be processed is incomplete, it means that the files to be processed have not been successfully inserted into the target database. If the file information of the files to be processed is completed, it means that the files to be processed have been successfully inserted into the target database. The processing module 64 extracts the splitting information of the original data file from the Redis database. If the splitting information is about the splitting breakpoints of the original data file, the splitting of the original data file continues based on the splitting breakpoints. If the splitting information is about the incomplete file information of the files to be processed, the files to be processed are re-inserted into the target database via a thread.
[0088] Optionally, the splitting module 61 is specifically used to split the original data file according to the splitting rules of the splitting key to obtain data fragments, wherein the splitting key corresponds to a data node;
[0089] The splitting module 61 is further used to add a corresponding current file write pointer to each obtained data slice, and write the data slice to the intermediate data file corresponding to the current file write pointer according to the file write pointer;
[0090] The splitting module 61 is further used to update the current file write pointer when the amount of data in the data shards written to the intermediate data file reaches a preset first threshold.
[0091] In this scenario example, the sharding key is typically one or more column fields. The sharding key is designed based on the principles of table and database sharding. This example uses the sharding key to determine which node each row of data in the original data file belongs to, so that each row of data is assigned to the corresponding data shard. The splitting module 61 can treat each data node as a data shard. Because the data size of each data node is inconsistent, when splitting the original data file according to the sharding key, a file write pointer is added to each resulting data shard. This write pointer points to the address to be written next. In this example, the file write pointer added to the data shard represents the intermediate data file to which the data shard should be written. The filename of the intermediate data file includes the corresponding node information. Specifically, the current file write pointer is added to each split data shard so that the data shard is written to the corresponding current intermediate data file. When the amount of data written reaches the rated size, the write pointer will be updated. For example, a first threshold for the amount of data written to the intermediate data file is preset. When the amount of data written to the intermediate data file reaches the first threshold, the file write pointer of the next data file is updated to the next intermediate data file so that the next data file is written to the next intermediate data file. In this example, after the amount of data written to the intermediate data file reaches the preset threshold, updating the file write pointer of the data file can prevent the intermediate data file from becoming too large.
[0092] Optionally, the splitting information can be the splitting breakpoint of the original data file. Since the splitting of the original data file is performed according to the sharding key rule, the splitting breakpoint of the original data file is updated in the Redis database every time the original data file is split.
[0093] On the other hand, the splitting information can be the file information of the files to be processed. Specifically, when a file to be processed is first split, its file information is recorded in the Redis database. At this time, the file information of the file to be processed is "incomplete," indicating that it has not yet been inserted into the target database. When multiple files to be processed are batch-inserted into the target database using the JDBC batch processing method, because the thread is processing asynchronously, after a file to be processed is successfully inserted into the target database, the Redis database is notified of the status information that the file to be processed has been successfully inserted into the target database, and the file information of the file to be processed is changed to "complete." Recording the file information of the files to be processed in the Redis database through asynchronous processing avoids preemption of the Redis database and ensures the accurate recording of splitting information.
[0094] Optionally, the splitting information for the original data file can be extracted from the Redis database first. If there are no splitting breakpoints for the original data file in the Redis database, it indicates that this is the first time the task of splitting the original data file and inserting it into the target database has been executed, or the splitting has been completed. Therefore, it can be further determined whether the splitting information includes the file information of the files to be processed. If the splitting information does not include the file information of the files to be processed, it indicates that the splitting of the original data file has not yet started. In this case, the splitting work of the original data file can be performed directly, and the splitting breakpoints for the original data file recorded in the Redis database are updated based on the splitting progress. Then, each split file to be processed is inserted into the target database using JDBC multi-threading, and the file information of the files to be processed recorded in the Redis database is updated. If the splitting information includes the file information of the files to be processed, the work of inserting the files marked as incomplete in the file information into the target database continues, and the file information of the files to be processed recorded in the Redis database is updated. If the Redis database records a split breakpoint for the original initial file, then the splitting of the original data file continues based on the recorded split breakpoint.
[0095] Optionally, after each file to be processed, derived from the original data file, has been inserted into the target database via a thread, it indicates that all work on the original data file has been completed. Therefore, it is unnecessary to record any splitting information about the original data file in the Redis database, and the splitting information recorded in the Redis database can be cleared. At this point, all files to be processed have been successfully processed, and the intermediate data files generated during the process will also be cleaned up.
[0096] Optionally, when inserting the files to be processed into the target database using the JDBC batch processing method, JDBC connection resources need to be enabled. Since multiple threads exist in the thread pool, the JDBC connection resources can be enabled only when the number of files to be processed reaches the second threshold, thus improving the overall processing efficiency of the files. After the JDBC connection resources are enabled, the files to be processed are asynchronously inserted into the target database using the JDBC batch processing method with multiple threads. Optionally, the second threshold can be selected as the number of threads in the thread pool, ensuring that each file to be processed has a corresponding thread for processing.
[0097] The splitting module splits the original data file into multiple intermediate data files, and then further splits each intermediate data file into multiple files to be processed according to a preset data size. The insertion module inserts each file to be processed into the target database asynchronously through multiple threads, with each thread processing one file. The splitting information of the original data files is recorded in a Redis database. When a breakpoint retrieval is required, the processing module retrieves the splitting information of the original data files from the Redis database and continues the subsequent processing of the original data files based on this information. Based on the method provided in this embodiment, by processing the multiple files to be processed asynchronously through multiple threads and recording the splitting information of the original data files in a Redis database, there is no contention for the Redis database, ensuring accurate recording of splitting information. Furthermore, when the splitting information of a data file needs to be retrieved to continue processing, it can be retrieved from the Redis database in a timely manner, greatly improving the efficiency of breakpoint retrieval and importing the data files into the database.
[0098] The apparatus provided in this embodiment can be used to execute the method embodiment shown above. Its implementation principle and technical effect are similar, and will not be described again here.
[0099] Example 3
[0100] Figure 7 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application, as shown below. Figure 7 As shown, the electronic device includes:
[0101] The system includes a processor 291 and a memory 292 communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory. It may also include a communication interface 293 and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via the bus 294. The communication interface 293 can be used for information transmission. The processor 291 can invoke logical instructions stored in the memory 292 to execute the method described in Embodiment 1.
[0102] Furthermore, the logic instructions in the aforementioned memory 292 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0103] The memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, implementing the method of Embodiment 1 described above.
[0104] The memory 292 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 292 may include high-speed random access memory and may also include non-volatile memory.
[0105] This application provides a non-transitory computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods described in the foregoing embodiments.
[0106] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention filed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0107] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for re-extracting breakpoints, characterized in that, The method includes: The original data file is split into multiple intermediate data files, and each intermediate data file is further split into multiple files to be processed according to a preset data size. By using multiple threads to process asynchronously, each file to be processed is inserted into the target database, with each thread processing one file to be processed. The splitting information of the original data file is recorded in the Redis database, including: after each file to be processed is split, the file information of the split file is recorded in the Redis database; when the thread starts importing the data of the file to be processed into the target database, feedback is sent to the Redis database so that the Redis database updates the file information of the file to be processed to indicate that it is complete; the splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split file to be processed. The system retrieves the splitting information of the original data file from the Redis database and continues to process the original data file based on the splitting information. The subsequent processing includes: determining whether there is a splitting breakpoint for the original data file in the Redis database; if there is, continuing to split the original data file based on the splitting breakpoint; if not, determining whether there is file information for the file to be processed in the Redis database; if there is, continuing to process the file to be processed through multi-threading based on the file information for the file to be processed; if not, directly splitting the original data file.
2. The method according to claim 1, characterized in that, The process of splitting the original data file to obtain multiple intermediate data files includes: The original data file is split according to the splitting rules of the splitting key to obtain data fragments, and the splitting key corresponds to a data node; For each data slice obtained, add a corresponding current file write pointer, and write the data slice to the intermediate data file corresponding to the current file write pointer according to the file write pointer; When the amount of data in the data fragments written to the intermediate data file reaches a preset first threshold, the current file write pointer is updated.
3. The method according to claim 1, characterized in that, The method of inserting each file to be processed into the target database through asynchronous processing using multiple threads includes: It is determined whether the number of files to be processed has reached a preset second threshold. If the number of files to be processed has reached the second threshold, each file to be processed is inserted into the target database asynchronously through multiple threads using JDBC batch processing.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: Once all the file information for all pending files related to the original data file in the splitting information recorded in the Redis database has been completed, clear the splitting information for the original data file recorded in the Redis database.
5. A breakpoint re-lifting device, characterized in that, The device includes: The splitting module is used to split the original data file into multiple intermediate data files, and then split each intermediate data file into multiple files to be processed according to a preset data size. The insert module is used to insert each file to be processed into the target database in an asynchronous manner using multiple threads, with each thread processing one file to be processed. The recording module is used to record the splitting information of the original data file in a Redis database, including: recording the file information of the split files in the Redis database after each file to be processed is completed; when a thread starts importing the data of the file to be processed into the target database, it sends a feedback to the Redis database so that the Redis database updates the file information of the file to be processed to indicate completion; the splitting information of the original data file includes the splitting breakpoint of the original data file and the file information of the split files. The processing module is used to obtain the splitting information of the original data file from the Redis database, and to continue the subsequent processing of the original data file based on the splitting information. The subsequent processing includes: determining whether there is a splitting breakpoint for the original data file in the Redis database; if so, continuing to split the original data file based on the splitting breakpoint; if not, determining whether there is file information for the file to be processed in the Redis database; if so, continuing to process the file to be processed through multi-threading based on the file information; if not, directly splitting the original data file.
6. The apparatus according to claim 5, characterized in that, The splitting module is specifically used to split the original data file according to the splitting rules of the splitting key to obtain data fragments, wherein the splitting key corresponds to a data node; The splitting module is further configured to add a corresponding current file write pointer to each obtained data slice, and write the data slice to the intermediate data file corresponding to the current file write pointer according to the file write pointer; The splitting module is further configured to update the current file write pointer when the amount of data in the data shards written to the intermediate data file reaches a preset first threshold.
7. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-4.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-4.
Citation Information
Patent Citations
File processing method and file processing device based on sub-libraries and sub-statements
CN107402950A
File splitting method and device
CN111625505A