Data loading method, device, electronic device and storage medium
By parallel processing based on data blocks and parsing termination locations during data loading, the problem of low loading performance of a single large file in the prior art is solved, and more efficient data loading and memory management are achieved.
Patent Information
- Application Number
- CN202510038928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-10
AI Technical Summary
When facing the loading of a single large file, the performance of existing data loading technology is degraded and cannot fully utilize the central processor resources. The memory usage is high, and frequent application and release leads to performance impact.
Based on the data block space occupied parameters, the data file is read using the reading thread to generate the initial data block, and based on the parsing termination position, the parsing thread analyzes the target data block to generate the initial row data. This method adopts blocked data block queue and data row queue, avoids data replication and makes full use of the central processor resources.
Effectively utilize system processing capabilities, reduce memory usage, avoid frequent memory application and release, and improve data loading speed and performance.
Smart Images

Figure CN119440670B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic digital data processing, and in particular to a data loading method, device, electronic equipment and storage medium. Background Art
[0002] In the process of database data loading (importing) and data conversion, it is necessary to read, parse (splitting rows / columns), load or convert the data in the data file. The existing data loading technology reads files in row order and cannot read multiple rows in batches. When facing the loading of a single large file, the performance will be significantly reduced; after reading the data in the file, the processing process will often copy the data to generate new data objects, such as reading row data and parsing column data, resulting in frequent memory application and release, thus affecting performance. Therefore, how to make full use of CPU resources, reduce memory usage, and improve performance when facing the loading of a single large file is a technical problem that needs to be solved urgently. Summary of the invention
[0003] In view of the above problems, the present invention provides a data loading method, device, electronic device and storage medium.
[0004] According to a first aspect of the present invention, a data loading method is provided, comprising: based on a data block occupied space parameter, using a reading thread to read a data file to generate an initial data block; processing the initial data block to generate a target data block, wherein the target data block represents a data block marked with a parsing termination position, and the target data block is stored in a blocking data block queue; based on the parsing termination position, using a parsing thread to parse the target data block to generate at least one initial row of data, wherein the parsing thread is parallel to the reading thread, the initial row of data represents data marked with a column position, and at least one initial row of data is stored in a blocking data row queue; based on the column position, performing column data assignment processing on at least one initial row of data in the blocking data row queue to obtain target row data; and using a data manipulation language to batch load the target row of data into a target database.
[0005] Optionally, the initial data block is processed to generate a target data block, including: for an i-th initial data block among the M initial data blocks, when it is determined that the first parsing termination position in the i-1-th target data block is the non-block tail position of the i-1-th target data block, based on a forward and downward rule, determining the position after the first parsing termination position as the first block header position of the i-th target data block, wherein 1≤i≤M-1, and M is a positive integer ≥2; performing a reverse search on the i-th initial data block to determine a first target line break character, wherein the reverse search is Searching from the block tail position of the initial data block to the block header position of the initial data block, the first target line break character represents the first line break character found in the reverse direction in the i-th initial data block; marking the position of the first target line break character as the first parsing termination position in the i-th target data block; determining the block tail position of the i-th initial data block as the first block tail position of the i-th target data block; generating the i-th target data block according to the first block header position of the i-th target data block, the first parsing termination position of the i-th target data block and the first block tail position of the i-th target data block.
[0006] Optionally, processing the initial data block to generate the target data block also includes: determining the block header position of the ith initial data block as the second block header position of the ith target data block when determining that the first parsing termination position in the i-1th target data block is the block tail position of the i-1th target data block; performing a reverse search on the ith initial data block to determine a second target line break character; marking the position of the second target line break character as the second parsing termination position in the ith target data block, wherein the second target line break character represents the first line break character found by reverse search in the ith initial data block; determining the block tail position of the ith initial data block as the second block tail position of the ith target data block; and generating the ith target data block according to the second block header position of the ith target data block, the second parsing termination position of the ith target data block, and the second block tail position of the ith target data block.
[0007] Optionally, based on the parsing termination position, a parsing thread is used to parse a target data block in a blocking data block queue to generate at least one initial row of data, including: using multiple parsing threads to parse multiple target data blocks in a blocking data block queue in parallel; for each target data block, the bytes in the target data block are read based on a byte reading order; when a column separator is read, the position of the column separator is marked as a column position; when a line break is read, the position of the line break is marked as a row position; when the parsing termination position is read, the parsing termination position is marked as a row position, and parsing of the target data block is stopped; and at least one initial row of data is determined based on the row position and the column position.
[0008] Optionally, based on the column position, column data assignment processing is performed on at least one initial row data in the blocking data row queue to obtain target row data, including: using multiple reading threads to read multiple initial row data from the blocking data row queue in parallel; for each initial row data, calling the data reading class to obtain multiple column data based on the column position; performing column data assignment processing on the multiple column data to obtain the target row data.
[0009] Optionally, batch loading into the target database includes: based on a batch loading row number parameter, batch loading multiple target row data into the target database in parallel using multiple loading threads.
[0010] Optionally, the data loading method also includes: when the reading thread reads the tail byte in the data file, generating an empty data block, wherein the number of empty structure block data is the same as the number of parsing threads, the empty structure block data represents a data block with a termination read flag, and the empty data block is stored in a blocking data block queue; when it is determined that the parsing thread reads an empty data block from the blocking data block queue, executing exit program processing for the parsing thread.
[0011] The second aspect of the present invention provides a data loading device, including: a reading module, which is used to read a data file using a reading thread based on a data block occupied space parameter to generate an initial data block; a processing module, which is used to process the initial data block to generate a target data block, wherein the target data block represents a data block marked with a parsing termination position, and the target data block is stored in a blocking data block queue; a parsing module, which is used to parse the target data block using a parsing thread based on the parsing termination position to generate at least one initial row of data, wherein the parsing thread is parallel to the reading thread, the initial row of data represents data marked with a column position, and at least one initial row of data is stored in a blocking data row queue; an assignment module, which is used to perform column data assignment processing on at least one initial row of data in the blocking data row queue based on the column position to obtain the target row data; and a loading module, which is used to batch load the target row data into a target database using a data manipulation language.
[0012] The third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned data loading method.
[0013] The fourth aspect of the present invention further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned data loading method.
[0014] According to the data loading method, device, electronic device and storage medium provided by the present invention, the initial data block is generated by reading the data file using a reading thread, and then the target data block is generated; based on the parsing termination position, the target data block is parsed using a parsing thread to generate at least one initial row data, and the parsing thread is parallel to the reading thread; based on the column position, at least one initial row data in the blocking data row queue is assigned column data to obtain the target row data, so as to load the target row data into the target database in batches. Since the initial data block is read from the data file based on the data block occupied space parameter, the system processing capacity is effectively utilized; the reference method of the parsing termination position is used in the process of generating the target data block to calculate the position and record the position, without copying the data, avoiding the frequent application and release of the memory; in addition, by setting the parsing thread and the reading thread in parallel, it is possible to realize the reading of the initial data block to be read and the parsing of the target data block to be parsed at the same time, thereby making full use of the central processing unit CPU resources, improving the data loading speed and improving the performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.
[0016] Figure 1 The application scenario of the data loading method and device according to the embodiment of the present invention is shown.
[0017] Figure 2 A flow chart of a data loading method according to an embodiment of the present invention is shown.
[0018] Figure 3 A flowchart of generating initial row data according to an embodiment of the present invention is shown.
[0019] Figure 4 A structural block diagram of a data loading device according to an embodiment of the present invention is shown.
[0020] Figure 5 A block diagram of an electronic device suitable for implementing a data loading method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0021] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.
[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0023] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0024] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0025] In the process of implementing the present invention, it is found that the existing data loading technology reads files in row order and cannot read multiple rows in batches. When facing the loading of a single large file, the performance will be significantly reduced; after reading the data in the file, the processing process will often copy the data to generate new copies, such as row data and column data, resulting in frequent application and release of memory, thereby affecting performance. With the rapid development of large-scale distributed databases, more data can be quickly swallowed and can grow linearly with the increase of nodes, so the performance of existing data loading technology can no longer meet the needs of technological development, especially when data is only stored in one file, it cannot be processed in parallel and cannot fully utilize CPU resources, resulting in a sharp decline in performance.
[0026] In view of this, an embodiment of the present invention provides a data loading method, device, electronic device and storage medium. The method includes: based on the data block occupied space parameter, using a reading thread to read a data file to generate an initial data block; processing the initial data block to generate a target data block, the target data block represents a data block marked with a parsing termination position, and the target data block is stored in a blocking data block queue; based on the parsing termination position, using a parsing thread to parse the target data block to generate at least one initial row of data, the parsing thread is parallel to the reading thread, the initial row of data represents data marked with a column position, and at least one initial row of data is stored in a blocking data row queue; based on the column position, performing column data assignment processing on at least one initial row of data in the blocking data row queue to obtain the target row of data; using a data manipulation language to batch load the target row of data into a target database.
[0027] In the technical solution of the present invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0028] Figure 1 The application scenario of the data loading method and device according to the embodiment of the present invention is shown.
[0029] like Figure 1 As shown, the business system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0030] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0032] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The cluster distributed database is deployed on the cluster server 105. The server 105 includes multiple server nodes that can analyze and process the received user request and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to the user request) to the terminal device.
[0033] It should be noted that the data loading method provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the data loading device provided in the embodiment of the present invention can generally be set in the server 105. The data loading method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the data loading device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0034] Alternatively, the data loading method provided in the embodiment of the present invention may also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the data loading device provided in the embodiment of the present invention may also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may be provided in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0036] It should be noted that the sequence numbers of the operations in the following method are only used as representations of the operations for the purpose of description, and should not be regarded as representing the execution order of the operations. Unless explicitly stated, the method does not need to be executed completely in the order shown.
[0037] Figure 2 A flow chart of a data loading method according to an embodiment of the present invention is shown.
[0038] like Figure 2 As shown, the method 200 includes operations S210 to S250.
[0039] In operation S210, based on the space occupied by the data block parameter, a data file is read by using a reading thread to generate an initial data block.
[0040] Optionally, before reading the data file, use the Config class to load configuration parameters from the configuration file, use loadFromFile(String filename) to load functions from the configuration file, and then use the loading class FastLoader to start the running framework and implement overall scheduling.
[0041] Optional configuration parameters include source parameters, target parameters, parsing parameters, optimization parameters, etc. Source parameters include data file storage path parameters dataFile, data file name parameters, etc. Target parameters include database connection parameters jdbcDbUrl, database target table information parameters tableName, etc. Parsing parameters include line break lineSplit (supports multi-byte), column separator colSplit (supports multi-byte); Optimization parameters include data block space parameter blockSize, data block queue depth parameter blockQueueDepth, data block parsing parallelism parameter dataParserThreads, data row queue depth parameter rowQueueDepth, data loading parallelism parameter dbLoadThreads, batch loading row number parameter rowsPerBatch, etc.
[0042] Optionally, after starting the running framework, the queue initialization class BlockedQueue is used to initialize the blocking data block queue based on the data block queue depth parameter, and the blocking data row queue is initialized based on the data row queue depth parameter.
[0043] Optionally, the data block space parameter represents the size of the initial data block space. The data block space parameter can be set to an integer multiple of the input / output (I / O) block size. If the row is a fixed length, the data block space parameter can also be set to an integer multiple of the row length, which can improve reading speed and performance.
[0044] Optionally, after the queue is initialized, the run() function in the reading class BlockReaderService is called to start multiple reading threads, and the BlockReader subclass in the BlockReaderService class is used to complete the parallel reading of data files by multiple reading threads, reading 100M of data each time to generate an initial data block of 100M.
[0045] In operation S220, the initial data block is processed to generate a target data block.
[0046] Optionally, the block queue depth parameter is used to limit the memory usage and central processing unit (CPU) usage of the blocking block queue, and no CPU consumption is incurred during the blocking period.
[0047] Optionally, while reading the data file to generate the initial data block, the initial data block can be processed based on a reverse search algorithm, and the end position of the last complete line in the initial data block can be marked as the parsing termination position to obtain a target data block with a data structure.
[0048] For example, the space occupied by a data block parameter is 100M. The space occupied by a data block parameter generally includes many lines. To improve the reading performance, 100M initial data blocks are continuously read from the data file. Each time an initial data block is read, it is necessary to determine whether the end of the initial data block is the end of a complete line. If it is not the end of a complete line, it is necessary to calculate the end position of the last complete line in the initial data block and mark it as the parsing end position, and append the remaining half line of data in the initial data block to the next initial data block to obtain the target data block corresponding to the next initial data block. If the first line of the current initial data block is a complete line and the last line is not a complete line, the size of the target data block corresponding to the current data block is still 100M, and the size of the target data block corresponding to the next initial data block is the sum of the data size of the last line of the incomplete line of the current data block and 100M. If the end of the initial data block is the end of a complete line, no appending is required to reduce memory operations. The parsing end position is used, and no memory copy is involved.
[0049] Optionally, the target data block can be stored in a blocking data block queue. When storing the target data block in a blocking data block queue, the depth of the blocking data block queue must be checked first. When the blocking data block queue does not reach the depth set by the block queue depth parameter blockQueueDepth, it means that the memory is not full, and it will return immediately to start reading the next initial data block. Therefore, when the parsing capability and data loading capability are sufficient, the file will continue to be read, and the input / output (I / O) maximum capability of the system can be reached. When the depth set by the block queue depth parameter is reached, it means that the memory is full, the parsing capability or the data loading capability is limited, and the target data block will be blocked from being put in. No CPU consumption will be generated during the blocking period to save CPU usage.
[0050] In operation S230 , based on the parsing termination position, the target data block is parsed using a parsing thread to generate at least one initial row of data.
[0051] Optionally, call the run() function in the parsing class BlockParserService to start multiple parsing threads, use the BlockParser subclass in the BlockParserService class to complete the parallel reading of multiple target data blocks in the blocking data block queue by multiple parsing threads, parse from the block header position of the target data block to the parsing end position based on the line break character and column separator, and generate at least one initial row of data. The number of entries in the initial row of data is the same as the number of line breaks parsed from the block header position to the parsing end position.
[0052] Optionally, the parsing end position is used to identify the position where the parsing of the data in the target data block stops.
[0053] Optionally, the column position is determined based on the column separator, and at least one initial row of data marked with a column position data structure is sequentially placed into a blocking data row queue.
[0054] Optionally, the initial row data can be stored in a blocking data row queue. When storing the initial row data in a blocking data row queue, the depth of the blocking data row queue needs to be checked first. When the blocking data row queue does not reach the depth set by the row queue depth parameter rowQueueDepth, it means that the memory is not full, and it will return immediately to start putting the next initial row data; when it reaches the depth set by the row queue depth parameter, it means that the memory is full, and the initial row data will be blocked. No CPU consumption will be generated during the blocking period.
[0055] Optionally, the i-th target data block obtained from the i-th initial data block is placed in a blocking data block queue. While the reading thread reads the i+1-th initial data block, the parsing thread can parse the i-th target data block in the blocking data block queue. When there are multiple parsing threads, the multiple parsing threads are responsible for parsing the first i target data blocks that have been placed in the blocking data block queue in parallel, and the reading thread is responsible for continuing to read the initial data block from the data file. The parsing thread is parallel to the reading thread, and there is no need to wait until all the data in the data file is read before starting parsing, which fully utilizes the capabilities of the input and output (IO) devices, improves the parallel processing capabilities of the entire process, and fully utilizes the central processing unit resources.
[0056] In operation S240 , based on the column position, column data assignment processing is performed on at least one initial row data in the blocking data row queue to obtain target row data.
[0057] Optionally, call the run() function in the loading class DBLoader to start multiple loading threads, use the DataRow subclass in the DBLoader class to complete the reading of each column value in the initial row data in the blocking data row queue based on the column position, assign each column value to the attribute variable, and obtain the processed target row data.
[0058] In operation S250 , the target row data is batch loaded into the target database using a data manipulation language.
[0059] The target row data is batch loaded into the target database based on the loading thread using the Structured Query Language (SQL).
[0060] Optionally, for scenarios where the above method is used for data conversion and loading, you can add custom operations based on the acquired data and position according to your needs, such as when calculating the entire data block, parsing the data block into rows and columns, or before entering the queue or after leaving the queue, so that users can perform custom extensions based on the above method.
[0061] Optionally, since the initial data block is read from the data file based on the data block space occupation parameter, the system processing capacity is effectively utilized; the reference method of the parsing termination position in the target data block generation process is to calculate the position and record the position, without copying the data, thus avoiding the frequent application and release of memory; in addition, by setting the parsing thread and the reading thread in parallel, it is possible to read the initial data block to be read and parse the target data block to be parsed at the same time, thereby making full use of the central processing unit CPU resources, increasing the data loading speed and improving the performance.
[0062] Optionally, both the initial data block and the target data block include M; wherein the initial data block is processed to generate the target data block, including: for the i-th initial data block among the M initial data blocks, when it is determined that the first parsing termination position in the i-1-th target data block is the non-block tail position of the i-1-th target data block, based on the forward and downward rule, the position after the first parsing termination position is determined as the first block header position of the i-th target data block, wherein 1≤i≤M-1, and M is a positive integer ≥2; reverse search is performed on the i-th initial data block to determine the first target line break position Characters, wherein the reverse search is to search from the block tail position of the initial data block to the block header position of the initial data block, the first target line break character represents the first line break character found in the reverse search in the i-th initial data block; the position of the first target line break character is marked as the first parsing termination position in the i-th target data block; the block tail position of the i-th initial data block is determined as the first block tail position of the i-th target data block; and the i-th target data block is generated according to the first block header position of the i-th target data block, the first parsing termination position of the i-th target data block and the first block tail position of the i-th target data block.
[0063] For example, if the data block space parameter is 100M, 100M data is read from the start byte in the data file to get the 0th initial data block, and a reverse search is performed from the block tail position of the 0th initial data block to the block head position to determine the first target line break, and the position of the first target line break is marked as the first parsing end position in the 0th target data block. The 0th target data block has the same block head position and block tail position as the 0th initial data block.
[0064] For example, taking the above-mentioned 0th initial data block as an example, 100M data is read from the last byte of the 0th initial data block to obtain the 1st initial data block, and the block header position of the 1st initial data block is the position of the last byte of the 0th initial data block.
[0065] Optionally, the initial data block read based on the setting of the data block space parameter generally contains many lines, which can improve the reading performance. However, the end position of the initial data block read may not be the end of a line, which will result in the need to parse incomplete lines during parsing. Therefore, based on the forward merge rule and the reference method of the parsing end position, it can be ensured that only complete lines in the target data block are parsed.
[0066] Optionally, the reading thread reads the i-th initial data block and determines whether the first parsing termination position in the i-1-th target data block is the block end position of the i-1-th target data block. When the first parsing termination position is a non-block end position of the i-1-th target data block, it represents that the tail row of the i-1-th target data block is an incomplete first half row data, and the incomplete first half row data of the i-1 target data block needs to be merged downward into the i-th initial data block. Thus, the i-th initial data block contains the incomplete first half row data and the incomplete second half row data of this row, forming a complete first row. In this process, the i-th first parsing termination position can be synchronously marked in the i-th initial data block. Thus, after completing the marking and first row filling, the i-th target data block is obtained. When the first parsing end position is the end of the block of the i-1th target data block, it means that there is no incomplete row data in the i-1th target data block. At this time, no additional processing is performed to ensure that the first line in the i-th target data block generated based on the i-th initial data block is complete row data.
[0067] For example, when it is determined that the first parsing end position in the 0th target data block is the non-block end position of the 0th target data block, it means that there is incomplete row data in the 0th target data block, and based on the forward and downward rule, the position after the first parsing end position is determined as the first block header position of the 1st target data block.
[0068] Optionally, the first parsing termination position is the position of the line break of the last complete line in the target data block, and the position after the first parsing termination position is the line header position of the incomplete line data.
[0069] For example, a reverse search is performed from the block tail position of the first initial data block to the block head position to determine the first target line break character, and the position of the first target line break character is marked as the first parsing termination position in the first target data block. The first block head position of the first target data block is different from the block head position of the first initial data block, and the first block tail position of the first target data block is the same as the block tail position.
[0070] Optionally, the data in the first target data block is data from the first block header position of the first target data block to the first block tail position of the first target data block, and the first target data block represents a data block having a data structure marked with a first parsing termination position.
[0071] Optionally, the first parsing termination position is determined based on the reverse search algorithm, ensuring that the last line of data parsed when the target data block is subsequently parsed is a complete line of data; based on the forward merge rule, the first line of data parsed when the target data block is subsequently parsed is a complete line of data, which ensures to the greatest extent that only complete lines of data in the target data block are parsed without missing data in the data file. The calculation speed of reverse search and forward merge is fast, the CPU consumption is low, and the data file can be read quickly and continuously.
[0072] Optionally, processing the initial data block to generate the target data block also includes: determining the block header position of the ith initial data block as the second block header position of the ith target data block when determining that the first parsing termination position in the i-1th target data block is the block tail position of the i-1th target data block; performing a reverse search on the ith initial data block to determine a second target line break character; marking the position of the second target line break character as the second parsing termination position in the ith target data block, wherein the second target line break character represents the first line break character found by reverse search in the ith initial data block; determining the block tail position of the ith initial data block as the second block tail position of the ith target data block; and generating the ith target data block according to the second block header position of the ith target data block, the second parsing termination position of the ith target data block, and the second block tail position of the ith target data block.
[0073] Optionally, the second analysis end position and the first analysis end position are both analysis end positions, and there is no specific substantive distinction between the second analysis end position and the first analysis end position.
[0074] Optionally, the first block header position and the second block header position are both block header positions, and there is no specific substantive distinction between the first block header position and the second block header position.
[0075] Optionally, the first block tail position and the second block tail position are both block tail positions, and there is no specific substantive distinction between the first block tail position and the second block tail position.
[0076] Optionally, when there is no incomplete row data in the previous target data block, the data of the first row of the current initial data block is complete row data, and there is no need to use the forward merge rule.
[0077] For example, taking the above-mentioned 0th initial data block as an example, when the first parsing end position in the 0th target data block is determined to be the block end position of the 0th target data block, it means that there is no incomplete row data in the 0th target data block, and the block header position of the 1st initial data block is directly determined as the second block header position of the 1st target data block.
[0078] For example, a reverse search is performed from the block tail position of the first initial data block to the block head position to determine the first target line break character, and the position of the first target line break character is marked as the second parsing termination position in the first target data block. The second block head position of the first target data block is the same as the block head position of the first initial data block, and the second block tail position of the first target data block is the same as the block tail position of the first initial data block.
[0079] Optionally, the data in the first target data block is data from the second block header position of the first target data block to the second block tail position of the first target data block, and the first target data block represents a data block having a data structure marked with a second parsing termination position.
[0080] Optionally, M target data blocks are generated based on the above method, which will not be described in detail here.
[0081] Figure 3 A flowchart of generating initial row data according to an embodiment of the present invention is shown.
[0082] like Figure 3 As shown, generating at least one initial row of data includes operations S310 to S360.
[0083] In operation S310 , multiple parsing threads are used to parse multiple target data blocks in a blocking data block queue in parallel.
[0084] In operation S320 , for each target data block, bytes in the target data block are read based on a byte reading order.
[0085] In operation S330, when a column delimiter is read, a position of the column delimiter is marked as a column position.
[0086] In operation S340, when a line break character is read, the position of the line break character is marked as a line position.
[0087] In operation S350, when the parsing end position is read, the parsing end position is marked as a row position, and the parsing of the target data block is stopped.
[0088] In operation S360, at least one initial row data is determined according to the row position and the column position.
[0089] Optionally, the BlockParser subclass in the BlockParserService class is used to complete the parallel reading of multiple target data blocks in the blocking data block queue by multiple parsing threads, and one parsing thread is corresponding to the parsing of one target data block.
[0090] Optionally, the byte reading order may be to read bytes in sequence from the block header position to the block tail position of the target data block.
[0091] Optionally, for each target data block, the parsing will split the lines according to the specified line break character and split the columns according to the specified column separator character. During parsing, the bytes in the target data block are read based on the byte reading order. When the column separator character is read, the position of the column separator is marked as the column position; when the line break character is read, an object instance is generated and the row position is marked; when the parsing end position is read, the parsing end position is marked as the row position, and parsing is stopped, and the target data block is parsed.
[0092] Optionally, there is one row position mark, and one initial row data is obtained by parsing; there are multiple row position marks, and multiple initial row data are obtained by parsing. Each initial row data is row data of a structure marked with a column position.
[0093] Optionally, during the parsing process, a parsing thread can be implemented to parse a target data block at a time based on the parseRows() function of the BlockParser class. After the entire target data block is parsed, at least one initial row of data is sequentially placed into the blocking data row queue.
[0094] Optionally, during the parsing process, based on the hasNextRow() function and the getNextRow() function, when parsing a target data block, an initial row of data can be immediately put into the blocking data row queue without waiting for the entire target data block to be parsed, so as to control the parsing granularity more finely.
[0095] Optionally, after one traversal, the entire target data block can be parsed without any redundant operations. The reference method of the parsing end position does not involve memory copying and only traverses once, which greatly improves the computing performance and saves memory usage.
[0096] Optionally, based on the column position, column data assignment processing is performed on at least one initial row data in the blocking data row queue to obtain target row data, including: using multiple reading threads to read multiple initial row data from the blocking data row queue in parallel; for each initial row data, calling the data reading class to obtain multiple column data based on the column position; performing column data assignment processing on the multiple column data to obtain the target row data.
[0097] Optionally, the data reading class DataRow in the DBLoader class is used to complete reading of each column value in the initial row data in the blocking data row queue based on the column position to obtain multiple column data.
[0098] Optionally, the target table information in the target database is determined based on the database connection parameter jdbcDbUrl, the target table information parameter tableName in the database, etc. There are multiple column attribute variables in the target table information. The data of each column is assigned to the corresponding column attribute variable using the data query language (Structured Query Language, SQL) to obtain the processed target row data.
[0099] Optionally, batch loading into the target database includes: based on a batch loading row number parameter, batch loading multiple target row data into the target database in parallel using multiple loading threads.
[0100] Optionally, the batch loading row number parameter rowsPerBatch is the row number parameter for batch submission to the target database. When the specified batch row number rowsPerBatch is reached, the loading thread batch updates the target table in the target database.
[0101] Optionally, multiple loading threads are started based on the data loading parallelism parameter, and one loading thread submits a batch of target row data. When all target row data corresponding to the target data block are confirmed to be loaded into the target database, or when there is no other use, the target data block is cleaned up in time to release the block data resources.
[0102] Optionally, parallel processing of data loading improves loading performance, greatly reduces the data loading cycle, and provides a good user experience. In addition, the target data block is always kept as a copy from the data file to the cleanup. During this life cycle, reverse search calculation, forward and downward calculation, and analytical calculation do not generate duplicate data, which improves processing performance while effectively saving memory space.
[0103] Optionally, the data loading method also includes: when the reading thread reads the tail byte in the data file, generating an empty data block, wherein the number of empty structure block data is the same as the number of parsing threads, the empty structure block data represents a data block with a termination read flag, and the empty data block is stored in a blocking data block queue; when it is determined that the parsing thread reads an empty data block from the blocking data block queue, executing exit program processing for the parsing thread.
[0104] Optionally, when the reading thread reads the tail byte in the data file, it means that the data in the data file has been read, and an empty data block is generated. The empty data block has no data but has a structure with a read termination mark.
[0105] Optionally, the termination read flag is used to notify the parsing thread that has read this empty data block to exit the parsing program. For example, the termination read flag may be an isEnd: true flag.
[0106] Optionally, multiple empty data blocks are put into a blocking data block queue. When the parsing thread reads an empty data block, it means that there are no target data blocks to be parsed in the blocking data block queue, and the parsing thread is exited.
[0107] Optionally, the number of empty structure block data is the same as the number of parsing threads to ensure that after all parsing threads have finished parsing the current target data block, there is no target data block to be parsed in the blocking data block queue. At this time, all parsing threads can read the empty data block in parallel. For example, if the number of parsing threads is 3, 3 empty data blocks are generated to ensure that the 3 parsing threads can read the empty data block in parallel and exit the program respectively.
[0108] Optionally, the setting of the structure and number of empty data blocks ensures that each parsing thread can read an empty data block when there is no target data block to be parsed, and then exit the program without paying attention to whether other parsing threads have completed parsing, and without the need for overall scheduling. The parsing capability increases linearly with the number of CPU cores, effectively saving CPU resources.
[0109] Based on the above data loading method, the present invention also provides a data loading device. Figure 4 The device is described in detail.
[0110] Figure 4 A structural block diagram of a data loading device according to an embodiment of the present invention is shown.
[0111] like Figure 4 As shown, the data loading device 400 of this embodiment includes a reading module 410 , a processing module 420 , a parsing module 430 , an assignment module 440 and a loading module 450 .
[0112] The reading module 410 is used to read the data file using a reading thread based on the data block occupied space parameter to generate an initial data block. In one embodiment, the reading module 410 can be used to perform the operation S210 described above, which will not be described in detail here.
[0113] The processing module 420 is used to process the initial data block to generate a target data block, wherein the target data block represents a data block marked with a parsing termination position, and the target data block is stored in a blocking data block queue. In one embodiment, the processing module 420 can be used to perform the operation S220 described above, which will not be repeated here.
[0114] The parsing module 430 is used to parse the target data block based on the parsing termination position using a parsing thread to generate at least one initial row of data, wherein the parsing thread is parallel to the reading thread, the initial row of data represents data marked with a column position, and the at least one initial row of data is stored in a blocking data row queue. In one embodiment, the parsing module 430 can be used to perform the operation S230 described above, which will not be described in detail here.
[0115] The assignment module 440 is used to perform column data assignment processing on at least one initial row of data in the blocking data row queue based on the column position to obtain target row data. In one embodiment, the assignment module 440 can be used to perform the operation S240 described above, which will not be repeated here.
[0116] The loading module 450 is used to load the target row data into the target database in batches using a data operation language. In one embodiment, the loading module 450 can be used to perform the operation S250 described above, which will not be described in detail here.
[0117] Optionally, the processing module 420 includes a first processing submodule, a second processing submodule, a third processing submodule, a fourth processing submodule and a fifth processing submodule.
[0118] The first processing submodule is used to determine, for an i-th initial data block among the M initial data blocks, a position subsequent to the first parsing end position as a first block header position of the i-th target data block based on a forward merge rule when determining that the first parsing end position in the i-1-th target data block is a non-block tail position of the i-1-th target data block, wherein 1≤i≤M-1, and M is a positive integer ≥2.
[0119] The second processing submodule is used to perform a reverse search on the i-th initial data block to determine a first target line break character, wherein the reverse search is to search from the end position of the initial data block to the head position of the initial data block, and the first target line break character represents the first line break character found by reverse search in the i-th initial data block.
[0120] The third processing submodule is used to mark the position of the first target line break as the first parsing termination position in the i-th target data block.
[0121] The fourth processing submodule is used to determine the block tail position of the i-th initial data block as the first block tail position of the i-th target data block.
[0122] The fifth processing submodule is used to generate an i-th target data block according to the first block header position of the i-th target data block, the first parsing termination position of the i-th target data block and the first block tail position of the i-th target data block.
[0123] Optionally, the processing module 420 further includes a sixth processing sub-module, a seventh processing sub-module, an eighth processing sub-module, a ninth processing sub-module and a tenth processing sub-module.
[0124] The sixth processing submodule is used to determine the block header position of the i-th initial data block as the second block header position of the i-th target data block when determining that the first parsing end position in the i-1-th target data block is the block tail position of the i-1-th target data block.
[0125] The seventh processing submodule is used to perform a reverse search on the i-th initial data block to determine a second target line break character.
[0126] The eighth processing submodule is used to mark the position of the second target line break as the second parsing termination position in the i-th target data block, wherein the second target line break represents the first line break found in the i-th initial data block by reverse search.
[0127] The ninth processing submodule is used to determine the block tail position of the i-th initial data block as the second block tail position of the i-th target data block.
[0128] The tenth processing submodule is used to generate an i-th target data block according to the second block header position of the i-th target data block, the second parsing termination position of the i-th target data block and the second block tail position of the i-th target data block.
[0129] Optionally, the parsing module 430 includes a first parsing submodule, a second parsing submodule, a third parsing submodule, a fourth parsing submodule, a fifth parsing submodule and a sixth parsing submodule.
[0130] The first parsing submodule is used to parse multiple target data blocks in the blocking data block queue in parallel using multiple parsing threads.
[0131] The second parsing submodule is used to read bytes in each target data block based on a byte reading order.
[0132] The third parsing submodule is used to mark the position of the column separator as the column position when the column separator is read.
[0133] The fourth parsing submodule is used to mark the position of the line break character as the line position when a line break character is read.
[0134] The fifth parsing submodule is used to mark the parsing end position as a row position when the parsing end position is read, and stop parsing the target data block.
[0135] The sixth parsing submodule is used to determine at least one initial row data according to the row position and the column position.
[0136] Optionally, the loading module 450 includes a first loading submodule, a second loading submodule and a third loading submodule.
[0137] The first loading submodule is used to read a plurality of initial row data from the blocking data row queue in parallel by using a plurality of reading threads.
[0138] The second loading submodule is used to call the data reading class for each initial row of data and obtain multiple column data based on the column position.
[0139] The third loading submodule is used to perform column data assignment processing on multiple column data to obtain target row data.
[0140] Optionally, any multiple modules among the reading module 410, the processing module 420, the parsing module 430, the assignment module 440 and the loading module 450 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. Optionally, at least one of the reading module 410, the processing module 420, the parsing module 430, the assignment module 440 and the loading module 450 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a suitable combination of any of them. Alternatively, at least one of the reading module 410 , the processing module 420 , the parsing module 430 , the assignment module 440 , and the loading module 450 may be at least partially implemented as a computer program module, and when the computer program module is executed, a corresponding function may be performed.
[0141] Figure 5 A block diagram of an electronic device suitable for implementing a data loading method according to an embodiment of the present invention is shown.
[0142] Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0143] like Figure 5As shown, the computer electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 to the random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include an on-board memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0144] In RAM 503, various programs and data required for the operation of electronic device 500 are stored. Processor 501, ROM 502 and RAM 503 are connected to each other via bus 504. Processor 501 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 502 and / or RAM 503. It should be noted that the program can also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in one or more memories.
[0145] Optionally, the electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage portion 508 as needed.
[0146] Optionally, the method flow according to the embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-mentioned functions defined in the system of the embodiment of the present invention are executed. Optionally, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0147] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the data loading method according to the embodiment of the present invention is implemented.
[0148] Optionally, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include, but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.
[0149] For example, optionally, the computer-readable storage medium may include the ROM 502 and / or the RAM 503 described above and / or one or more memories other than the ROM 502 and the RAM 503 .
[0150] An embodiment of the present invention also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the data loading method provided by the embodiment of the present invention.
[0151] When the computer program is executed by the processor 501, the above functions defined in the system / device of the embodiment of the present invention are performed. Optionally, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0152] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 509, and / or installed from the removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0153] Optionally, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages, specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. It can be understood by those skilled in the art that the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways, even if such a combination or combination is not explicitly recorded in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recorded in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0155] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A data loading method, characterized in that: The method comprises: Based on the space occupied by the data block parameter, the data file is read using the reading thread to generate the initial data block; Processing the initial data block to generate a target data block includes: For an i-th initial data block among the M initial data blocks, when it is determined that the first parsing end position in the i-1-th target data block is the non-block tail position of the i-1-th target data block, a position after the first parsing end position is determined as the first block header position of the i-th target data block based on a forward and downward rule, wherein 1≤i≤M-1, and M is a positive integer ≥2; Performing a reverse search on the i-th initial data block to determine a first target line break character, wherein the reverse search is to search from the end of the initial data block to the beginning of the initial data block, and the first target line break character represents the first line break character found by reverse search in the i-th initial data block; Marking the position of the first target line break as the first parsing termination position in the i-th target data block; Determine the block tail position of the i-th initial data block as the first block tail position of the i-th target data block; Generate an i-th target data block according to a first block header position of the i-th target data block, a first parsing termination position of the i-th target data block, and a first block tail position of the i-th target data block, wherein the target data block represents a data block marked with a parsing termination position, the target data block is stored in a blocking data block queue, and the initial data block and the target data block each include M; Based on the parsing termination position, the target data block is parsed using a parsing thread to generate at least one initial row of data, wherein the parsing thread is parallel to the reading thread, the initial row of data represents data marked with column positions, and the at least one initial row of data is stored in a blocking data row queue; Based on the column position, performing column data assignment processing on at least one of the initial row data in the blocking data row queue to obtain target row data; and The target row data is batch loaded into the target database using a data manipulation language.
2. The method according to claim 1, characterized in that The processing of the initial data block to generate a target data block further comprises: In the case where it is determined that the first parsing termination position in the i-1th target data block is the block tail position of the i-1th target data block, determining the block header position of the i-th initial data block as the second block header position of the i-th target data block; Performing a reverse search on the i-th initial data block to determine a second target line break character; Marking the position of the second target line break as the second parsing termination position in the i-th target data block, wherein the second target line break represents the first line break found in the i-th initial data block by reverse search; Determine the block tail position of the i-th initial data block as the second block tail position of the i-th target data block; The i-th target data block is generated according to the second block header position of the i-th target data block, the second parsing termination position of the i-th target data block and the second block tail position of the i-th target data block.
3. The method according to claim 1, characterized in that The blocking data block queue stores a plurality of target data blocks; Wherein, based on the parsing termination position, using a parsing thread to parse the target data block in the blocking data block queue to generate at least one initial row of data includes: Utilizing a plurality of the parsing threads to parse a plurality of the target data blocks in the blocking data block queue in parallel; For each of the target data blocks, bytes in the target data block are read based on a byte reading order; When a column separator is read, marking the position of the column separator as the column position; When a line break character is read, the position of the line break character is marked as the line position; When the parsing end position is read, the parsing end position is marked as the row position, and the parsing of the target data block is stopped; At least one of the initial row data is determined according to the row position and the column position.
4. The method according to claim 1, characterized in that: The performing column data assignment processing on at least one of the initial row data in the blocking data row queue based on the column position to obtain the target row data comprises: Using a plurality of the reading threads to read a plurality of the initial row data in parallel from the blocking data row queue; For each of the initial row data, calling a data reading class, and obtaining a plurality of column data based on the column position; Perform column data assignment processing on the plurality of column data to obtain the target row data.
5. The method according to claim 1, characterized in that: The bulk loading into the target database includes: Based on the batch loading row number parameter, multiple target row data are batch loaded into the target database in parallel using multiple loading threads.
6. The method according to claim 1, characterized in that The method further comprises: When the reading thread reads the tail byte in the data file, an empty data block is generated, wherein the number of the empty data blocks is the same as the number of the parsing threads, the empty data block represents a data block with a termination read flag, and the empty data block is stored in the blocking data block queue; When it is determined that the parsing thread has read the empty data block from the blocking data block queue, an exit program process is executed on the parsing thread.
7. A data loading device, characterized in that: The device comprises: A reading module is used to read the data file using a reading thread based on the space occupied by the data block to generate an initial data block; a processing module, used to process the initial data block to generate a target data block, the processing module comprising a first processing submodule, a second processing submodule, a third processing submodule, a fourth processing submodule and a fifth processing submodule, A first processing submodule is configured to, for an i-th initial data block among the M initial data blocks, determine, based on a forward and downward rule, a position after the first parsing end position as a first block header position of the i-th target data block when it is determined that the first parsing end position in the i-1-th target data block is a non-block tail position of the i-1-th target data block, wherein 1≤i≤M-1, and M is a positive integer ≥2; A second processing submodule is configured to perform a reverse search on the i-th initial data block to determine a first target line break character, wherein the reverse search is to search from the end of the initial data block to the beginning of the initial data block, and the first target line break character represents the first line break character found by reverse search in the i-th initial data block; A third processing submodule, configured to mark the position of the first target line break as the first parsing termination position in the i-th target data block; A fourth processing submodule, configured to determine the block tail position of the i-th initial data block as the first block tail position of the i-th target data block; a fifth processing submodule, configured to generate an i-th target data block according to a first block header position of the i-th target data block, a first parsing termination position of the i-th target data block, and a first block tail position of the i-th target data block, wherein the target data block represents a data block marked with a parsing termination position, and the target data block is stored in a blocking data block queue; A parsing module, configured to parse the target data block using a parsing thread based on the parsing termination position to generate at least one initial row of data, wherein the parsing thread is in parallel with the reading thread, the initial row of data represents data marked with column positions, the at least one initial row of data is stored in a blocking data row queue, and the initial data block and the target data block each include M; an assignment module, configured to perform column data assignment processing on at least one of the initial row data in the blocking data row queue based on the column position to obtain target row data; and The loading module is used to load the target row data into the target database in batches using a data operation language.
8. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more instructions, Wherein, when the one or more instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Million-level excel data quick and stable import system
CN110275918A
Data batch import method and device, electronic equipment and storage medium
CN117827979A