File reading and writing method and device, equipment, storage medium and program product

By using the format conversion interface in Gaussian database to convert the row data storage format file into columnar data storage format file and reverse conversion, the problem that Gaussian database cannot directly import and export Parquet files is solved, achieving efficient data import and export and reducing resource consumption and cost.

CN120216459APending Publication Date: 2025-06-27CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510270155.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Gaussian databases cannot directly import or export Parquet files, resulting in high resource consumption and high cost of integrating into ETL processes.

Method used

The row data storage format file is divided through the preset first format conversion interface, and the data block of the preset size is obtained and written into the pre-created column data storage format file; the data in the column data storage format file is parsed and converted through the preset second format conversion interface, and the row data storage format data is obtained and written into the original row data storage format file.

Benefits of technology

It realizes efficient import and export of Parquet files by Gaussian database, reduces resource consumption and reduces the cost of integrating into ETL processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216459A_ABST
    Figure CN120216459A_ABST
Patent Text Reader

Abstract

The invention discloses a file reading and writing method and device, equipment, a storage medium and a program product, and relates to the technical field of data processing, and the method comprises the following steps: obtaining a pipeline / disk line data storage format file; segmenting the row data file into data blocks with a preset size through a preset first format conversion interface, and writing the data blocks with the preset size into a pre-created column data storage format file; and performing analysis and format conversion on the pre-created column type data storage format file through a preset second format conversion interface to obtain row type data storage format data, and writing the row type data storage format data into the pipeline / disk row type data storage format file. According to the method, a function of importing a column type data storage format file into a Gaussian database or exporting a data table and landing the data table into the column type data storage format file is realized in a form of a plug-in file format conversion thread.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a file reading and writing method, apparatus, device, storage medium, and program product. Background Art

[0002] The GaussDB series of databases are domestic self-developed databases produced by Huawei, including OLTP (Online Transaction Processing) and OLAP (Online Analytical Processing) products, with characteristics such as being distributed, having high performance, high availability, and being easy to expand. Since its release, it has been successively adopted by major domestic banking enterprises to build the underlying financial technology for the next generation and replace the old imported commercial databases. Currently, the Gauss series of databases are widely used by multiple banking enterprises and have a leading market share and extensive influence in the domestic banking industry.

[0003] Currently, the Gauss database provides the GDS (Gauss Data Service) server as a basic file import and export tool, which does not support reading and writing Parquet (columnar data storage format) files, and cannot directly import Parquet files into the Gauss database or export Gauss data tables to Parquet files. Generally, tools such as Spark and Hive are used to convert between Parquet files and row-based table files, but they consume a large amount of resources and usually require a distributed environment; moreover, the cost of integrating them into the existing ETL process is relatively high. Summary of the Invention

[0004] The main purpose of this application is to provide a file reading and writing method, apparatus, device, storage medium, and program product, aiming to solve the technical problem of how to efficiently import and export Parquet files in the Gauss database.

[0005] To achieve the above object, this application proposes a file reading and writing method, and the file reading and writing method includes:

[0006] Obtain a pipeline / disk row-based data storage format file;

[0007] Split the data in the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and write the data blocks of the preset size into a pre-created columnar data storage format file;

[0008] Parse and convert the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0009] In one embodiment, the step of splitting the data in the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and writing the data blocks of the preset size into the pre-created columnar data storage format file includes:

[0010] Use a row record delimiter to split the data in the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size corresponding to the pipeline / disk row-based data storage format file, and put the data blocks of the preset size corresponding to the pipeline / disk row-based data storage format file into a consumption queue;

[0011] Consume the data blocks of the preset size in parallel in the consumption queue, parse the fields in the data blocks of the preset size, and write the parsed fields into multiple pre-created columnar data storage format files by column.

[0012] In one embodiment, the step of using a row record delimiter to split the data in the pipeline / disk row-based data storage format file to obtain data blocks of a preset size corresponding to the pipeline / disk row-based data storage format file, and putting the data blocks of the preset size corresponding to the pipeline / disk row-based data storage format file into a consumption queue further includes:

[0013] Allocate the data in the pipeline row-based data storage format file to a buffer, and cyclically read the data in the pipeline row-based data storage format file in the buffer until the buffer is full;

[0014] Match the row record delimiter to the data in the pipeline row-based data storage format file in reverse order from the end position of the buffer, and use the data in the pipeline row-based data storage format file before the first row record delimiter as a data block of the preset size;

[0015] According to the position and length of the data block of the preset size, confirm whether the data block of the preset size is at the end position in the buffer;

[0016] If so, put the data block of the preset size into the consumption queue.

[0017] In one embodiment, the step of splitting the data in the pipeline / disk row-based data storage format file by using the row record delimiter to obtain data blocks of a preset size corresponding to the pipeline / disk row-based data storage format file, and putting the data blocks of the preset size corresponding to the pipeline / disk row-based data storage format file into the consumption queue includes:

[0018] Offset the data in the disk row-based data storage format file to a specified position through a preset file random read interface;

[0019] Read the data in the disk row-based data storage format file from the specified position, and match the data in the disk row-based data storage format file with the row record delimiter in ascending order. The data in the disk row-based data storage format file from the starting position of the specified position to the first row record delimiter and in between is used as one data block of the preset size;

[0020] Confirm whether the data block of the preset size is at the end position according to the position and length of the data block of the preset size;

[0021] If so, put the data block of the preset size into the consumption queue.

[0022] In one embodiment, the step of consuming the data blocks of the preset size in parallel in the consumption queue, parsing the fields in the data blocks of the preset size, and writing the parsed fields into multiple pre-created columnar data storage format files by columns includes:

[0023] Obtain the data block from the consumption queue, and parse the data block to obtain the specified fields and the starting position and length of the specified fields;

[0024] Use an array of starting positions to record the starting positions of the specified fields, and use an array of lengths to record the lengths of the specified fields;

[0025] Traverse the data block according to the order of the specified fields and the classification delimiter to fill the array of starting positions and the array of lengths;

[0026] Convert the data blocks in the array of starting positions and the data blocks in the array of lengths into columnar data storage format respectively to obtain data in columnar data storage format;

[0027] Write the data in columnar data storage format into multiple pre-created columnar data storage format files.

[0028] In one embodiment, the step of parsing and format-converting the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row-based data storage format data and writing the row-based data storage format data into the pipeline / disk row-based data storage format file includes:

[0029] Parsing each of the pre-created columnar data storage format files through a preset second format conversion interface to obtain the data in columnar data storage format;

[0030] Reading the data in columnar data storage format in multiple threads and converting the data in columnar data storage format into row-based data storage format to obtain the row-based data storage format data;

[0031] Writing the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0032] In one embodiment, the step of reading the data in columnar data storage format in multiple threads and converting the data in columnar data storage format into row-based data storage format to obtain the row-based data storage format data includes:

[0033] Reading the data in columnar data storage format of each thread in sequence and assembling the read data in columnar data storage format into row-based data storage format data using a specified record separator and a row record separator.

[0034] In addition, to achieve the above object, the present application also provides a file reading and writing device, which includes:

[0035] An acquisition module, configured to acquire a pipeline / disk row-based data storage format file;

[0036] A first data processing module, configured to split the data in the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size and write the data blocks of the preset size into a pre-created columnar data storage format file;

[0037] A second data processing module, configured to parse and format-convert the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row-based data storage format data and write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0038] In addition, to achieve the above object, the present application also provides a file reading and writing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the file reading and writing method as described above.

[0039] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the file reading and writing method described above are implemented.

[0040] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the file reading and writing method described above are implemented.

[0041] The embodiments of the present application provide a file reading and writing method, device, equipment, storage medium, and program product. The method includes obtaining a pipeline / disk row-based data storage format file; splitting the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and writing the data blocks of the preset size into a pre-created column-based data storage format file; parsing and performing format conversion on the data in the pre-created column-based data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and writing the row-based data storage format data into the pipeline / disk row-based data storage format file. The method realizes the function of importing a column-based data storage format file into a Gaussian database or exporting a data table and landing it as a column-based data storage format file in the form of an external file format conversion thread. During the whole process, data is directly interacted through a pipeline without generating temporary files, thereby enhancing the flexibility and efficiency of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the file reading and writing method of the present application;

[0045] Figure 2 It is a schematic flowchart provided for field parsing and Parquet writing in Embodiment 1 of the present application;

[0046] Figure 3 It is a schematic flowchart provided for Parquet parsing and DAT generation in Embodiment 1 of the present application;

[0047] Figure 4 It is a schematic flowchart of the file reading and writing method provided in the first embodiment of the present application;

[0048] Figure 5 It is a schematic overall flowchart of DAT to Parquet provided for the file reading and writing method of the present application;

[0049] Figure 6 It is a schematic flowchart of the reading and splitting of DAT files provided for the file reading and writing method of the present application;

[0050] Figure 7 It is a schematic flowchart of the reading and splitting of disk DAT files provided for the file reading and writing method of the present application;

[0051] Figure 8 It is a schematic overall flowchart of Parquet to DAT provided for the file reading and writing method of the present application;

[0052] Figure 9 It is a schematic diagram of the module structure of the file reading and writing device in the embodiment of the present application;

[0053] Figure 10 It is a schematic diagram of the device structure of the hardware operating environment involved in the file reading and writing method in the embodiment of the present application.

[0054] The realization of the purpose, functional features and advantages of the present application will be further described in combination with the embodiments and with reference to the accompanying drawings. Detailed implementation manners

[0055] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0056] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.

[0057] The main solution of the embodiment of the present application is: obtaining a pipeline / disk row-based data storage format file; reading and splitting the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and writing the data blocks of the preset size into a pre-created columnar data storage format file; parsing and performing format conversion on the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and writing the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0058] In this embodiment, for the convenience of description, the personal computer is used as the execution subject for elaboration below.

[0059] In the current related field, the GaussDB series of domestic self-developed databases produced by Huawei is known, including OLTP and OLAP products, with characteristics such as being distributed, high-performance, highly available, and easy to expand. Since its release, it has been successively adopted by major domestic banks to build the financial technology foundation for the next generation and replace the old imported commercial databases. Currently, the Gauss series of databases has been widely adopted by major banks and has a leading market share and extensive influence in the domestic banking industry.

[0060] Parquet is a columnar data storage format, a top-level open-source project of the Apache Foundation, and uses the Apache License 2.0 protocol. In Parquet files, data is stored and scanned by column, and appropriate compression strategies are adopted according to the characteristics of each column's data, which can achieve a good compression ratio. Parquet is very suitable for OLAP scenarios. Compared with traditional row-based table files, it can greatly reduce storage space and IO bandwidth occupancy and accelerate file processing. Parquet is the mainstream storage format in the Hadoop ecosystem and has native support from mainstream data processing engines such as Hadoop, Spark, and Flink. It is one of the important basic file formats in the field of big data processing.

[0061] Among them, the Gauss database coexists with the Hadoop ecosystem. There is a need to frequently exchange a large amount of data between the two types of systems. In the relevant ETL process, the GDS server provided by the Gauss database is used as the basic data import and export tool. However, GDS only supports row-based table files (the file form can be disk files or pipes) and cannot directly support the reading and writing of Parquet files. Generally, tools such as Spark and Hive are used to convert between Parquet files and row-based table files, but the resource consumption is large, and a distributed environment is usually required; moreover, the cost of integrating into the existing ETL (Extract, Transform, Load) process is relatively high.

[0062] To achieve efficient import and export of Parquet files in a Gaussian database, this application provides a solution. It reads and splits a DAT file into multiple data blocks of a preset size using a pre-set interface, and writes these data blocks into a pre-created columnar data storage format Parquet file. Additionally, another pre-set interface is used to parse and convert the data in the Parquet file to generate data in DAT format, and rewrite this DAT format data back into a pipeline / disk DAT file. By defining a set of standardized interface specifications, this solution can not only flexibly expand functional interfaces based on different business requirements and file format characteristics, but also facilitate seamless integration into existing ETL (Extract, Transform, Load) processes, achieving efficient conversion between Parquet files and row-based table files in a single-machine mode, accelerating data processing speed, and improving system performance.

[0063] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, a personal computer will be taken as an example to illustrate this embodiment and the following embodiments.

[0064] Based on this, an embodiment of this application provides a file reading and writing method. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the file reading and writing method of this application.

[0065] In this embodiment, the file reading and writing method includes steps S10 to S30:

[0066] Step S10, obtain a pipeline / disk row-based data storage format file;

[0067] It should be noted that a disk row-based data storage format file refers to a data file stored on a computer hard disk or other types of storage media. A pipeline row-based data storage format file is a virtual file used for inter-process communication (IPC) in an operating system and belongs to a special file type.

[0068] A DAT file (row-based table file) refers to a file with a row-priority file storage format that customizes field delimiters and record delimiters. In the following content, row-based data storage format files will be collectively referred to as DAT files.

[0069] Step S20, split the data in the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and write the data blocks of the preset size into a pre-created columnar data storage format file;

[0070] It should be noted that the first format conversion interface set in this embodiment supports reading one or more pipeline / disk DAT files, and outputting one or more columnar data storage format files after conversion. In the following content, the columnar data storage format files are described as Parquet files. Among them, Parquet refers to the columnar data storage format specified in the Apache Parquet open source project.

[0071] It should also be noted that a data block of a preset size refers to splitting the entire file into multiple segments with a fixed or specified size when reading and processing a pipeline / disk DAT file. The size of these data blocks can be adjusted according to actual needs.

[0072] In this embodiment, for the implementation of the preset first format conversion interface, first connect the preset first format conversion interface to the pipeline / disk DAT file, then use the line record delimiter to split the data in the pipeline / disk DAT file, split the data into data blocks of appropriate size, and put the split data blocks into the consumption queue. When multiple consumers are started downstream, the data blocks in the queue are consumed in parallel, the fields are parsed, and finally written column by column into multiple pre-created Parquet files.

[0073] Next, step S20 will be described in detail. Step S20 may include steps S21 to S22:

[0074] Step S21, using the preset first format conversion interface to split the data in the pipeline / disk row-based data storage format file by the line record delimiter, obtaining data blocks of a preset size corresponding to the pipeline / disk row-based data storage format file, and putting the data blocks of a preset size corresponding to the pipeline / disk row-based data storage format file into the consumption queue;

[0075] It should be noted that the line record delimiter refers to a specific character or character sequence used to distinguish different records (lines) in a text file, including but not limited to line feed (\n), carriage return line feed (\r\n), carriage return (\r), and custom delimiters.

[0076] In a feasible embodiment, step S21 may include steps A211 to A214:

[0077] Step A211, allocate the data in the pipeline DAT file to the buffer, and loop to read the data in the pipeline DAT file in the buffer until the buffer is full;

[0078] It should be noted that a buffer refers to a section of memory area in computer science for temporarily storing data. As an intermediary in the data transmission process, it is used to balance the data processing speed differences between different devices or processes.

[0079] In this embodiment, for the reading and splitting of the pipeline DAT file, first, a memory area with a fixed size is created as the buffer. Data is continuously read from the pipeline DAT file and written into this buffer until the data in the buffer reaches the preset size or there is no more data to be read.

[0080] Through the above steps, by temporarily storing data in a fast-access memory area, the frequent read and write operations on external devices (such as hard disks, network interfaces, etc.) can be reduced, thereby accelerating the data processing speed.

[0081] Step A212, match the data in the pipeline row-based data storage format file with the row record delimiter in reverse order from the end position of the buffer, and the data in the pipeline row-based data storage format file before the first row record delimiter is used as a data block of the preset size;

[0082] Specifically, starting from the last byte of the buffer, scan in the starting direction to query the nearest row record delimiter. After finding the first row record delimiter, determine the position of the previous bit, and all the data before this position constitutes a complete data block of the preset size.

[0083] Through the above steps, ensure that the data block boundary is consistent with the logical record boundary, and avoid data blocks that cross records.

[0084] Step A213, check whether the data block of the preset size in the buffer meets the preset threshold;

[0085] Specifically, check whether the data block of the preset size found in the current buffer meets the preset threshold. By setting an appropriate threshold, the size of the data block can be controlled, neither too fragmented nor too large, so as to optimize the use of storage space and transmission efficiency.

[0086] Step A214, if so, put the data block of the preset size that meets the preset threshold into the consumption queue.

[0087] If it meets the requirements, put the data block of the preset size that meets the preset threshold into the consumption queue and wait for the next step of processing. If it does not meet the requirements, allocate and write the data block of the preset size that does not meet the preset threshold to the next buffer.

[0088] In the above steps, reading and splitting the pipeline DAT file through the preset first format conversion interface not only helps to ensure the consistency and accuracy of the data, but also can safely transfer it to the downstream processing module to ensure data integrity and processing efficiency.

[0089] In another feasible embodiment, step S21 may further include steps B211 to B214:

[0090] Step B211, offset the data in the disk row data storage format file to a specified position through the preset file random reading interface;

[0091] It should be noted that the file random reading interface refers to a function or method that allows a program to directly access data at any position in a file.

[0092] In addition to reading and splitting the pipeline row data storage format file, this embodiment also proposes a method for reading and splitting the disk DAT file.

[0093] First, use the preset file random reading interface to open the DAT file on the disk. And determine the starting point of the specified position to be read according to the business logic (such as the block offset start_offset recorded by the metadata). Then jump the DAT file pointer on the disk to the starting point of the specified position to prepare for subsequent reading.

[0094] Step B212, read the data in the disk row data storage format file from the specified position, and match the data in the disk row data storage format file with the line record separator in ascending order. The data in the disk row data storage format file from the starting position of the specified position to the first line record separator and in between is used as a data block of the preset size;

[0095] Specifically, read data of the preset size (such as 4KB) from the starting point of the specified position into the memory buffer. And scan byte by byte from the starting position of the buffer to find the line record separator (such as \n or \r\n). If the line record separator is found, the data from the starting position of the buffer to the first line record separator is used as a complete data block of the preset size; if not, expand the reading range until it is found or exceeds the preset threshold.

[0096] In addition, if the end boundary (such as end_offset) of the specified position has been read, it is directly truncated to the boundary.

[0097] Step B213, confirm whether the data block of the preset size is at the end position according to the position and length of the data block of the preset size;

[0098] Specifically, by recording the file handle, the position information and length information of the data block of the preset size are obtained, and the end position of the data block of the preset size in the sequence to which it belongs is calculated based on the position information and length information. Then, the end position of the data block of the preset size in the sequence to which it belongs is compared with the total length of the sequence to which it belongs. If the end position is equal to the total length of the sequence, it is confirmed that the data block of the preset size is the last data block of the preset size in the specified position; if it is not equal, it is confirmed that the data block of the preset size is not the last data block of the preset size in the specified position, and the read pointer is continuously offset with the data block size as the step length, and the data is re-read from the starting point of the specified position after the offset.

[0099] Step B214: If yes, then put the data block of the preset size into the consumption queue.

[0100] Specifically, after confirming that the data block of the preset size is the last data block of the preset size in the specified position, the data block of the preset size is put into the consumer queue to wait for the next step of processing.

[0101] Through the above steps, the system can efficiently extract formatted data blocks from the disk DAT file and safely pass them to the downstream processing module to ensure data integrity and processing efficiency.

[0102] Step S22, consuming the data blocks of the preset size in parallel in the consumption queue, parsing the fields in the data blocks of the preset size, and writing the parsed fields into a plurality of the pre-created columnar data storage format files by column.

[0103] It is worth noting that due to the conversion of row-based files into column-based files, a large number of memory misses will occur, causing performance bottlenecks. In order to increase the conversion rate and save computing resources, the cost of memory misses is borne during the DAT file reading stage. The processing of a single consumer thread is as follows: Figure 2 As shown, Figure 2 This is a flowchart of field parsing and Parquet writing.

[0104] Specifically, when multiple consumer work threads are started downstream, the consumers simultaneously obtain data blocks of a preset size from the consumer queue, wherein the data blocks of the preset size at least include original byte data and a start offset and an end offset.

[0105] Then, the data block of the preset size is parsed by field, for example, starting from the starting position of the data block of the preset size, scanning each character one by one, and recording the starting position and length of each field, wherein the length of the field is obtained by calculating the current position minus the field starting position.

[0106] Subsequently, store the parsed starting positions and lengths of the fields into the corresponding starting position array head[n][n] and length array len[n][n] respectively. Among them, the starting position array is used to store the byte offsets of each field in the data block; the length array is used to store the byte lengths of each field.

[0107] Next, use the KMP (Knuth-Morris-Pratt, string matching) algorithm to traverse the data block of a preset size according to the specified field delimiter and record delimiter to fill the starting position array head[n][n] and length array len[n][n]. Among them, the KMP algorithm is an improved string matching algorithm. It preprocesses the pattern string (i.e., the substring to be searched), constructs a partial match table (also known as the next array or failure function), and uses this table to perform efficient searches in the text string, avoiding unnecessary character comparisons; the specified field delimiter is used to separate the characters or strings of different fields in a record; the record delimiter is used to separate the characters or strings of different records.

[0108] When using the KMP algorithm to traverse each data block of a preset size, find the positions of all field delimiters and record delimiters. Whenever a delimiter is found, update the corresponding head[n][n] and len[n][n] arrays.

[0109] Using the obtained head[n][n] and len[n][n] arrays, extract the actual values of each field from the data block of a preset size. The actual value is obtained by intercepting the corresponding substring in the data block of a preset size according to the starting position and length of each field. Then, based on business requirements or predefined data patterns (Schema), determine the appropriate Parquet data type for each field. That is, determine the data type of each field in the Parquet file, such as INT32, DOUBLE, STRING, etc., so as to convert the original field values in the data block into the corresponding Parquet format. After completing the extraction of field values and the determination of types, use the API provided by the Parquet library to write these field values into the Parquet file according to the defined pattern (Schema).

[0110] In summary, use the preset first format conversion interface to efficiently convert the DAT file on the pipeline or disk into a Parquet file, thereby realizing the optimized import of the GaussDB (Gauss Database).

[0111] Step S30: Parse and perform format conversion on the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0112] In this embodiment, for the implementation of the preset second format conversion interface, first parse each Parquet file to obtain all rowgroups in the Parquet file. Then, start multi-threading at the rowgroup granularity, read and convert to the DAT format, and put it into the consumption queue. Finally, start one or more write threads as needed to write the data converted to the DAT format into the pipeline / disk DAT file.

[0113] It should be noted that since converting a columnar file to a row-based file will cause a large number of memory access misses, resulting in a performance bottleneck. To improve the conversion rate and save computing resources, bear the cost of memory access misses during the Parquet reading stage, and only read one piece of data each time and assemble it into DAT data. As Figure 3 shown, Figure 3 is a schematic diagram of the Parquet parsing and DAT generation process.

[0114] Next, a detailed description of step S30 is given. Step S30 may include steps S31 to S33:

[0115] Step S31: Parse each of the pre-created columnar data storage format files through a preset second format conversion interface to obtain the data in the columnar data storage format;

[0116] Specifically, parse the rowgroups in the Parquet file through a preset second format conversion interface to obtain the Parquet format data.

[0117] Step S32: Read the data in the columnar data storage format in a multi-threaded manner, and convert the data in the columnar data storage format to the DAT format to obtain the DAT format data;

[0118] Specifically, read the Parquet format data of each column in sequence, and only read one piece each time. Subsequently, use the specified record separator and line separator to assemble the read Parquet format data into a line of DAT record, that is, convert the Parquet format data to the DAT format to obtain the DAT format data.

[0119] Then, continuously write the DAT format data into the data buffer, and determine whether the number of DAT format data in the current data buffer reaches a preset threshold. If it does not reach the threshold, continue to process the remaining data in the rowgroup; if it reaches the threshold, put the DAT format data in the data buffer into the consumption queue for subsequent continuous processing.

[0120] Step S33, write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0121] When multiple consumer worker threads are started downstream, the consumer obtains the DAT format data from the consumption queue and writes it into the pipeline / disk DAT file.

[0122] In addition, it should be noted that in the above steps, the arrangement order of each piece of data in the Parquet rowgroup is the same as that of the original DAT, but the order between rowgroups is not guaranteed. If there is a need to maintain the original data sequence (for example, the original file is a sorted file), based on the custom metadata KeyValueMetadata specified in the Parquet file specification, when converting DAT to Parquet, record the position of the data block in the original DAT, starting from 0 and incrementing the number. Each data block corresponds to a rowgroup, and the rowgroup id and the DAT position form a key-value pair, which is recorded in the metadata.

[0123] When converting from Parquet format to DAT format, implement the ordered output of data blocks according to the above-recorded information. When parsing the rowgroup, sort it according to the DAT position serial number and pass it to the output thread. The output thread outputs in order to obtain a file with the same order as the original DAT.

[0124] In summary of the above steps, use the preset second format conversion interface to efficiently convert the Parquet file into a DAT file on the pipeline or disk, thereby realizing the optimized export of the GaussDB.

[0125] Through the method of the above embodiments, a pipeline / disk row-based data storage format file is obtained; the pipeline / disk row-based data storage format file is read and split through a preset first format conversion interface to obtain data blocks of a preset size, and the data blocks of the preset size are written into a pre-created columnar data storage format file; the data in the pre-created columnar data storage format file is parsed and format-converted through a preset second format conversion interface to obtain row-based data storage format data, and the row-based data storage format data is written into the pipeline / disk row-based data storage format file. This method realizes the function of importing a columnar data storage format file into a Gaussian database or exporting and landing a data table as a columnar data storage format file in the form of an external file format conversion thread. And throughout the process, data is directly interacted through a pipeline without generating temporary files, thereby enhancing the flexibility and efficiency of data processing.

[0126] Exemplarily, to help understand the implementation process of the file reading and writing method obtained by combining the above Embodiment 1, please refer to Figure 4 , Figure 4 which provides a schematic diagram of a brief process of a file reading and writing method. Specifically:

[0127] This embodiment proposes a set of file format conversion interface specifications (including a preset first format conversion interface and a preset second format conversion interface), as well as a transformation solution for the existing ETL process. This solution realizes the support for Parquet files in the form of an external file format conversion thread in the existing ETL, and data is interacted through a pipeline without landing temporary files.

[0128] In this embodiment, the conversion of a pipeline / disk DAT file to a Parquet file is realized through a preset first format conversion interface. As Figure 5 shown, Figure 5 is a schematic diagram of the overall process of converting DAT to Parquet. Specifically, it includes: connecting the preset first format conversion interface to the pipeline / disk DAT file -> splitting the data with a row record delimiter to split the data into data blocks of a preset size -> putting the preset size into a consumption queue -> multiple consumers can be started downstream to consume the data blocks in parallel, parse the fields, and write them into multiple Parquet files column by column.

[0129] Due to being limited by the pipeline sequential reading characteristic, therefore, one pipeline corresponds to one reading / splitting thread. As Figure 6 shown, Figure 6It is a schematic flow diagram for reading and splitting pipeline DAT files. When reading and splitting pipeline DAT files, first, open the pipeline DAT file, obtain the DAT-format data therein, and allocate this data to the buffer. Read cyclically in the buffer until the buffer is filled. Then search for the end of the buffer, match the line record delimiter for the data in the buffer in reverse order, and use the data up to and including the first line record delimiter as a complete data block. Next, determine whether the number of data blocks in the current buffer reaches the preset threshold. If so, put data blocks of the preset size into the consumption queue; if not, allocate and write the remaining data blocks to the next buffer.

[0130] As Figure 7 shown, Figure 7 It is a schematic flow diagram for reading and splitting disk DAT files. When reading and splitting disk DAT files, pre-split the data blocks, defer the actual reading process to the consumer thread to achieve parallel reading, and use the mmap technology to accelerate the reading. Specifically, first, open the disk DAT file and obtain the data blocks therein. According to the size of the data block, use the file random reading interface to jump the pointer of the data block to the specified position. Next, use the jumped specified position as the current position, start reading a small part of the DAT data from the current position, match the line record delimiter in sequence, and use the data from the starting position to before the first line record delimiter as a complete data block of the preset size. Subsequently, by calculating the starting position and length of the data block of the preset size, check whether the end position of the data block of the preset size is within the specified position. If so, confirm that the data block of the preset size is the last data block of the preset size in the specified position, and put the data block of the preset size into the consumer queue. If not, confirm that the data block of the preset size is not the last data block of the preset size in the specified position, then continue to offset the reading pointer by the size of the data block, and reread the data from the re-offset position.

[0131] After starting the downstream, obtain small data blocks of the preset size from the consumption queue, use the head[n][n] array to store the starting position of each field with the field as the parsing target, and use the len[n][n] array to store the length of each field. Subsequently, use the KMP algorithm to determine the delimiters between fields and the classification delimiter, and traverse the entire data block to ensure the correct identification and splitting of fields. Finally, according to the requirements of the Parquet file format, convert the DAT data of each column into the corresponding Parquet data type and construct the structure of the Parquet file.

[0132] In addition, this embodiment also implements the conversion of Parquet files into pipeline / disk DAT files through a preset second format conversion interface. As Figure 8 shown, Figure 8 FIG. 1 is a schematic diagram of the overall process of converting Parquet to DAT. First, a rowgroup is read and parsed from the Parquet file. After parsing the RowGroup, the Parquet-formatted data is extracted row by row, and only one row is extracted each time. Then, the read Parquet-formatted data is assembled into one DAT record using a specified record separator as the line separator. Subsequently, the DAT-formatted data is continuously written into the data buffer, and it is determined whether the number of DAT-formatted data in the current data buffer reaches a preset threshold, that is, to check whether there is still unprocessed data in the rowgroup. If there is still remaining unprocessed data in the current rowgroup, return to the step of "extracting Parquet-formatted data row by row" to continue processing. If all the data in the current rowgroup has been processed, the processed DAT-formatted data is put into the consumer queue. Finally, when multiple consumer worker threads are started downstream, the consumer obtains the DAT-formatted data from the consumption queue and writes it into the pipeline / disk DAT file.

[0133] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the file reading and writing method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0134] This application also provides a file reading and writing device. Please refer to Figure 9 , the file reading and writing device includes:

[0135] An acquisition module 10, configured to acquire a pipeline / disk row-based data storage format file;

[0136] A first data processing module 20, configured to split the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and write the data blocks of the preset size into a pre-created column-based data storage format file;

[0137] A second data processing module 30, configured to parse and perform format conversion on the data in the pre-created column-based data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0138] The file reading and writing device provided by this application adopts the file reading and writing method in the above embodiment, and can solve the technical problem of efficiently importing and exporting Parquet files in the Gaussian database. Compared with the prior art, the beneficial effects of the file reading and writing device provided by this application are the same as those of the file reading and writing method provided by the above embodiment, and other technical features in the file reading and writing device are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.

[0139] This application provides a file reading and writing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the file reading and writing method in the first embodiment above.

[0140] Refer to the following Figure 10 , which shows a schematic structural diagram of a file reading and writing device suitable for implementing the embodiments of this application. The file reading and writing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The file reading and writing device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0141] As Figure 10As shown, the file reading and writing device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the file reading and writing device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the file reading and writing device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a file reading and writing device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0142] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0143] The file reading and writing device provided by the present application adopts the file reading and writing method in the above-mentioned embodiment, and can solve the technical problem of efficiently importing and exporting Parquet files in a Gaussian database. Compared with the prior art, the beneficial effects of the file reading and writing device provided by the present application are the same as those of the file reading and writing method provided by the above-mentioned embodiment, and other technical features in the file reading and writing device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0144] It should be understood that each part disclosed in the present application may be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0145] As described above, this is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.

[0146] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the file reading and writing method in the above embodiments.

[0147] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0148] The above computer-readable storage medium can be included in the file reading and writing device; it can also exist separately and not be assembled into the file reading and writing device.

[0149] The above computer-readable storage medium carries one or more programs, which, when executed by a file reading and writing device, cause the file reading and writing device to: obtain a pipeline / disk row-based data storage format file; split the pipeline / disk row-based data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and write the data blocks of the preset size into a pre-created column-based data storage format file; parse and perform format conversion on the data in the pre-created column-based data storage format file through a preset second format conversion interface to obtain row-based data storage format data, and write the row-based data storage format data into the pipeline / disk row-based data storage format file.

[0150] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0152] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0153] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above file reading and writing method, which can solve the technical problem of efficiently importing and exporting Parquet files in the Gaussian database. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the file reading and writing method provided by the above embodiments, and will not be elaborated here.

[0154] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the file reading and writing method as described above are implemented.

[0155] The computer program product provided by the present application can solve the technical problem of efficiently importing and exporting Parquet files in the Gaussian database. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the file reading and writing method provided by the above embodiments, and will not be elaborated here.

[0156] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A file reading and writing method, characterized in that: The file reading and writing method comprises: Get pipeline / disk row data storage format file; Splitting the data in the pipeline / disk row data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and writing the data blocks of the preset size into a pre-created column data storage format file; The data in the pre-created columnar data storage format file is parsed and format-converted through a preset second format conversion interface to obtain row data storage format data, and the row data storage format data is written into the pipeline / disk row data storage format file.

2. The file reading and writing method according to claim 1, characterized in that: The step of dividing the data in the pipeline / disk row data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and writing the data blocks of the preset size into a pre-created column data storage format file comprises: The data in the pipeline / disk row data storage format file is divided by using a row record separator through a preset first format conversion interface to obtain data blocks of a preset size corresponding to the pipeline / disk row data storage format file, and the data blocks of the preset size corresponding to the pipeline / disk row data storage format file are placed into a consumption queue; The data blocks of the preset size are consumed in parallel in the consumption queue, fields in the data blocks of the preset size are parsed, and the parsed fields are written into a plurality of the pre-created columnar data storage format files by column.

3. The file reading and writing method according to claim 2, characterized in that: The step of using a row record separator to divide the data in the pipeline / disk row data storage format file to obtain a data block of a preset size corresponding to the pipeline / disk row data storage format file, and putting the data block of the preset size corresponding to the pipeline / disk row data storage format file into a consumption queue comprises: Allocating the data in the pipeline line data storage format file to a buffer, and cyclically reading the data in the pipeline line data storage format file in the buffer until the buffer is filled; Match the row record separator to the data in the pipeline line data storage format file in reverse order from the end position of the buffer, and use the data in the pipeline line data storage format file before the first row record separator as a data block of the preset size; Checking whether a data block of a preset size in the buffer satisfies a preset threshold; If so, a data block of a preset size that meets the preset threshold is placed in the consumption queue.

4. The file reading and writing method according to claim 2, characterized in that: The step of using a row record separator to divide the data in the pipeline / disk row data storage format file to obtain a data block of a preset size corresponding to the pipeline / disk row data storage format file, and putting the data block of the preset size corresponding to the pipeline / disk row data storage format file into a consumption queue also includes: Through a preset file random read interface, the data in the disk row data storage format file is offset to a specified position; Read the data in the disk row data storage format file from the specified position, and match the row record separators to the data in the disk row data storage format file in positive order, and use the data in the disk row data storage format file from the starting position of the specified position to the first row record separator and the data therebetween as a data block of the preset size; According to the position and length of the data block of the preset size, confirm whether the data block of the preset size is at the end position; If so, the data block of the preset size is placed in the consumption queue.

5. The file reading and writing method according to claim 2, characterized in that: The step of consuming the data blocks of the preset size in parallel in the consumption queue, parsing the fields in the data blocks of the preset size, and writing the parsed fields into a plurality of the pre-created columnar data storage format files by column comprises: Obtain the data block from the consumption queue, and parse the data block to obtain a specified field and a starting position and length of the specified field; Using a starting position array to record the starting position of the specified field, and using a length array to record the length of the specified field; Traversing the data block according to the specified field delimiter and record delimiter to fill the starting position array and the length array; Converting the data blocks in the starting position array and the data blocks in the length array into columnar data storage formats respectively to obtain data in columnar data storage formats; The data in the columnar data storage format is written into a plurality of pre-created columnar data storage format files.

6. The file reading and writing method according to claim 1, characterized in that: The step of parsing and formatting the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row data storage format data, and writing the row data storage format data into the pipeline / disk row data storage format file comprises: Parsing each of the pre-created columnar data storage format files through a preset second format conversion interface to obtain data in the columnar data storage format; Multithreading reads the data in the columnar data storage format, and converts the data in the columnar data storage format into a row data storage format to obtain the data in the row data storage format; The row data storage format data is written into the pipeline / disk row data storage format file.

7. The file reading and writing method according to claim 6, characterized in that: The multithreading reads the data in the columnar data storage format and converts the data in the columnar data storage format into a row data storage format. The step of obtaining the data in the row data storage format includes: The data in the columnar data storage format of each thread is read sequentially, and the data in the columnar data storage format read are assembled into data in the row data storage format using the specified record separator and row record separator.

8. A file reading and writing device, characterized in that: The file reading and writing device comprises: The acquisition module is used to obtain pipeline / disk row data storage format files; A first data processing module is used to split the data in the pipeline / disk row data storage format file through a preset first format conversion interface to obtain data blocks of a preset size, and write the data blocks of the preset size into a pre-created column data storage format file; The second data processing module is used to parse and convert the data in the pre-created columnar data storage format file through a preset second format conversion interface to obtain row data storage format data, and write the row data storage format data into the pipeline / disk row data storage format file.

9. A file reading and writing device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the file reading and writing method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the file reading and writing method according to any one of claims 1 to 7 are implemented.

11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the file reading and writing method according to any one of claims 1 to 7 are implemented.