Method and apparatus for high-performance analysis of a delimiter format file
Through the finite state deducer model and SIMD acceleration technology, the parallel processing bottleneck in the CSV file parsing process is solved, efficient parallel CSV file processing is achieved, and processing speed and CPU utilization is improved.
Patent Information
- Application Number
- CN202410300118.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-03-15
AI Technical Summary
There are parallel processing bottlenecks in the parsing and analysis of existing CSV files, resulting in performance losses and resource-intensive overhead, making it difficult to make full use of the hardware parallelism of multi-core systems.
The parallel scanning method based on the finite state deducer model is adopted to iteratively recognize control characters, generate bitmap indexes, and use SIMD accelerated delimiter recognition to realize parallel processing and query of files.
提升了CSV文件的处理速度和CPU利用率,减少了内存消耗,适应不同核心数量的环境,实现了高效的并行处理。
Smart Images

Figure CN118227669B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of information technology and data analysis, and particularly relates to a method and device for high-performance analysis of delimiter format files. Background Art
[0002] Parsing delimiter format files (Comma Separated Value, abbreviated as CSV) has been a long-standing challenging problem in the field of data analysis research. Its main task is to efficiently parse and analyze such delimiter format files and query the information required by users from the original files. Therefore, achieving fast parsing and real-time analysis is the focus of studying this problem. CSV files separated by delimiters, as a basic file format, increasingly appear in various applications. However, the inherent format of CSV files may hinder parallel processing, resulting in performance loss and significant overhead. Data extraction converts CSV files into a tabular structure, and common processing flows are often limited to block decoupling and handling nested delimiters, such as characters within fields or quotes, resulting in serial processing. The data extraction and conversion process is still a resource-intensive overhead and becomes a bottleneck in CSV data analysis. Parallel processing can improve the utilization rate of the CPU under a multi-core system and at the same time increase the speed of CSV data parsing and analysis.
[0003] CSV parallel processing technology is a high-performance method for CSV format data sets, which can greatly accelerate parsing and reduce costs. Different from the traditional idea of building an index for CSV files based on a serial thinking, parallel technology divides CSV files into equally sized blocks, builds indexes within each block respectively, and finally merges the indexes. However, the scalability and performance of existing research work still depend on the conversion of blocks and context verification. Faster and more complex hardware architectures highlight the CSV extraction and conversion bottlenecks in the entire analysis process, especially in parallel processing. To adapt to the exponential growth of computing throughput, manufacturers are working on expanding the number of cores and improving the Single Instruction Multiple Data (SIMD) capabilities. In order to fully utilize the potential of current hardware parallelism and benefit from the increasing number of cores, recent CSV algorithms and applications need to have scalability under large-scale hardware architectures. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method and device for high-performance analysis of delimiter format files.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for high-performance analysis of delimiter format files includes the following steps:
[0007] Sample the control characters in the input file iteratively to determine the symbol state and logical position of the control characters;
[0008] Determine the character conversion level to be selected in the finite state transducer model according to the control characters, and the character conversion levels include record level and field level;
[0009] Slice the input file into text blocks of equal size, put them into idle processing units, implement parallel scanning based on the finite state transducer model, and use SIMD to accelerate the recognition of delimiters to generate a bitmap index, which maps the logical positions of delimiters to physical positions;
[0010] Perform queries based on the bitmap index, including keyword search query mode and file union query mode.
[0011] Furthermore, the control characters include delimiters, double quotes, and escape characters. The delimiters include commas and line breaks. The part other than the control characters is called text characters. The symbol state refers to whether it exists in the sampling. The logical position refers to the relative position of the sampled control character in the text.
[0012] Furthermore, the iterative sampling of the control characters in the input file includes using a heuristic algorithm to check the control characters in the file. That is, as long as a certain control character in the control characters is encountered, it is considered that the file contains control characters and the check of this control character is stopped until all characters in the control characters are included or the end of the file is reached. And, use the SIMD instruction set to quickly judge whether a continuous string in the CSV file contains a certain type of control character, expanding the sampling range while maintaining the original sampling speed as much as possible.
[0013] Furthermore, the finite state transducer is a finite state transducer based on state minimization, including four states of IR, ER, IQ, and EQ and their corresponding conversions for CSV format file conversion, where IR represents inside the record, ER represents the end of the record, IQ represents inside the quotes, and EQ represents the end of the quotes.
[0014] Furthermore, determining the character conversion level to be selected in the finite state transducer model according to the control characters includes: when the sampling does not have double quotes, it is considered that the position of the delimiter can be determined through the character conversion at the record level; when the sampling has double quotes, it is considered that the field level needs to be combined for character conversion at the field level to determine the position of the delimiter.
[0015] Furthermore, process the bitmap index to simplify the query time, or directly perform queries without processing the bitmap index.
[0016] Further, in the keyword search query mode, the user provides the keywords to be queried, and the line numbers and column numbers of all exactly matching items found in the file are returned; in the file union query mode, the user provides the folder name to be queried and the keywords to be queried, then reads all CSV files in the folder directory, and returns the file names, line numbers, and column numbers of all exactly matching items found in the file.
[0017] A method and device for high-performance analysis of delimiter format files, including:
[0018] A control character sampling module, configured to sample control characters in an input file in an iterative manner to determine the symbol state and logical position of the control characters;
[0019] A conversion level selection module, configured to determine the character conversion level to be selected in a finite state transducer model according to the control characters, and the character conversion levels include a record level and a field level;
[0020] A bitmap index generation module, configured to divide an input file into text blocks of equal size, put them into idle processing units, implement parallel scanning based on the finite state transducer model, and use SIMD to accelerate the recognition of delimiters to generate a bitmap index, and the bitmap index maps the logical positions of delimiters to physical positions;
[0021] A query module, configured to perform queries based on the bitmap index, including a keyword search query mode and a file union query mode.
[0022] The advantages of the present invention compared with the prior art are as follows:
[0023] (1) Compared with other single-thread processing methods for tabular data, the present invention realizes parallel processing at the thread level and instruction level, and solves the speed bottleneck problem that CSV file processing is limited by the inherent format (multiple types of row and column delimiters) and can only be processed serially.
[0024] (2) The present invention establishes a hierarchical finite state transducer model, which can heuristically determine various CSV files containing different control characters, and select the simplest state for conversion therefrom, thereby improving the processing speed and reducing the memory consumption.
[0025] (3) The present invention has good scalability in different core number tests, can be adapted to be deployed in environments with different core numbers, and makes full use of the current environment to achieve parallelism. Description of the Drawings
[0026] Figure 1 It is the state of a CSV format string based on a finite state transducer model (FST).
[0027] Figure 2 It is a flowchart of a high-performance CSV processing method based on the FST model. Specific implementation manners
[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and drawings.
[0029] The present invention realizes a new CSV processing framework by exploring architecture-aware optimization and supporting various conditional query applications. This method designs a hierarchical finite-state transducer model (Finite-State Transducer, abbreviated as FST) with minimized state to achieve parallel CSV parsing and data conversion. This framework covers architecture-aware optimization, with parallel scanning and SIMD-based delimiter recognition to facilitate parallel CSV parsing. The present invention integrates the programming interface of the user's CSV query with the method of the present invention to achieve on-demand loading optimization, including keyword search and file union query programs.
[0030] The technical solution of the present invention determines the symbol state and logical position of the delimiter in each file based on iterative identification of symbols through heuristic control character sampling; designs a finite-state transducer (FST) model with minimized state, as Figure 1 shown. This model consists of character conversions at the record level and field level, and selects the appropriate level according to the control characters obtained in the sampling; divides the entire file into the same size and places it in each processing unit, and performs parsing operations through parallel scanning based on FST and SIMD-accelerated delimiter recognition; the most critical performance aspect stems from nested quotes, so the present invention sets efficient bitmap indexes for relative positions and global positions and maps the logical positions to physical positions; finally, designs a pattern query programming function for easy-to-use applications.
[0031] A method for high-performance analysis of a delimiter format file of the present invention proposes a high-parallelism CSV large file processing framework based on the FST model. The specific technical solution is as Figure 2 shown, including the following steps:
[0032] Step 1: Identify (sample) the control characters in the input file iteratively to determine the symbol state and logical position of the control characters;
[0033] Step 2: Determine the character conversion level to be selected in the FST model according to the control characters, and the character conversion level includes the record level and the field level;
[0034] Step 3: Split the input file into text chunks of equal size and place them in idle processing units. Implement parallel scanning based on FST, and use SIMD to accelerate the recognition of delimiters to generate a bitmap index, mapping the logical positions of delimiters to physical positions;
[0035] Step 4: Optionally, further process the bitmap index to simplify the query time, or directly perform a query without processing the bitmap index;
[0036] Step 5: Perform a query based on the constructed bitmap index, providing two query modes: keyword search and file union query.
[0037] Furthermore, the control characters in Step 1 include delimiters (comma, newline), double quotes, and escape characters. The part outside the control characters is usually called text characters, including all uppercase and lowercase letters, numbers from 0 to 9, common punctuation marks such as commas, periods, semicolons, colons (it should be noted that in CSV text, the content containing such characters often needs to be within double quotes), mathematical symbols such as plus sign (+), minus sign (-), equal sign (=), and whitespace characters (space characters). Text characters are the basic elements that make up words, sentences, and text content.
[0038] Furthermore, the symbol state in Step 1 refers to the existence during sampling, and the logical position refers to the relative position of the sampled control character in the text.
[0039] Furthermore, the iterative recognition symbols in Step 1 mainly consist of two parts: a heuristic algorithm and SIMD acceleration. The specific steps are as follows:
[0040] Step 1.1: The overall idea of the heuristic algorithm is to check the control characters in the file as quickly as possible. That is, as long as a certain control character in the control characters (such as double quotes) is encountered, it is considered that the file contains this control character and the check of this control character stops until all characters in the control characters are included or the end of the file is reached. At the same time, to ensure the sampling speed, the formula of the heuristic algorithm adopted in the present invention is as follows:
[0041]
[0042] k < log2(length - i)
[0043] where f(x) represents the position that the pointer should point to during each sampling, k represents the number of samplings, length represents the length of the text, i represents the position pointed to by the initial pointer, and the value of i is 100.
[0044] The remaining reasons for adopting this algorithm are as follows: in a CSV file, the data at the front end is often more complete and sufficient. Therefore, the algorithm of the present invention gradually reduces the sampling frequency as it moves towards the back end of the file; the initial value setting of i ensures that most of the file header can be skipped to reduce useless sampling.
[0045] Step 1.2: On the other hand, in order to maintain the trend of exponential growth in computing power, manufacturers have gradually expanded the number of cores and the single instruction multiple data (SIMD) capabilities. The hardware parallelism of a CPU containing multiple chip sets even exceeds that of a single chip and can be extended to multiple inherently parallel modules on a machine. Recent x86 CPUs can work on 256-bit SSE registers, that is, each register can hold 32 8-bit characters. The present invention focuses on CPU processors supporting SIMD instructions (vector instructions) for the parallel CSV extraction and analysis phase, but the method proposed by the present invention is also applicable to SIMD-based accelerators, including GPUs and FPGAs.
[0046] Therefore, the present invention can use the SIMD instruction set to process the comparison of 32-bit characters at one time, quickly judge whether a certain type of control character is included in the continuous string in the CSV file, and while expanding the sampling range, maintain the original sampling speed as much as possible.
[0047] Furthermore, the data set adopted in step 2 is specifically all CSV data sets obtained from the Kaggle data science repository. During the experiment of the present invention, CSV files with different structures were used, which spanned many categories, such as user IDs, e-commerce behaviors, competitions, educational materials, images, etc. The present invention used real-world CSV format data sets to evaluate the extraction and analysis performance. For the remaining microbenchmark evaluations, the present invention also used the data sets mentioned above. The sizes of these CSV files range from 328MB to 21GB, and all files use UTF-8 encoding.
[0048] Step 2 mainly consists of three parts: a finite state transducer (FST), a state transition mechanism, and state minimization. The specific steps are as follows:
[0049] Step 2.1: The present invention defines FST CSV =(S, T) as the complete character state set formed according to the file, where S is the preset state of the first character of each block, and T is the state transition relationship based on control characters and text characters. The complete model of the finite state transducer FST consists of all deterministic states ( Figure 1 the four states shown) converted from the CSV format file.
[0050] Step 2.2: The following attributes of CSV support the state transition mechanism: The number of initial character states (i.e., represented by a finite number of initial states) in the CSV formatted string is finite. All control characters in the format file separated by delimiters consist of delimiters (comma, line feed), double quotes, and escape characters.
[0051] According to the state transition of characters, the present invention summarizes four states of IR, ER, IQ, and EQ and their corresponding transitions to cover all possible situations. As Figure 1 shown, where IR refers to Inside Record, indicating inside the record; ER refers to End of Record, indicating the end of the record; IQ refers to Inside Quotation, indicating inside the quotation marks; EQ refers to End of Quotation, indicating the end of the quotation marks. Figure 1 The "*" in
[0052] refers to a wildcard character, that is, all symbols without additional indication except the delimiter.
[0053] The conversions corresponding to the above four states specifically include:
[0054] The transition from IR to ER: Specifically, when the current state is IR and the character passed at this time is a line feed, it is converted to ER; the transition from IR to IQ: Specifically, when the current state is IR and the character passed at this time is a double quote, it is converted to IQ; the transition from IR to IR: Specifically, when the current state is IR and the character passed at this time is other characters except the line feed and double quote, it is converted to IR.
[0054] The transition from ER to IQ: Specifically, when the current state is ER and the character passed at this time is a double quote, it is converted to IQ; the transition from ER to IR: Specifically, when the current state is ER and the symbol passed at this time is other characters, it is converted to IR.
[0055] The transition from IQ to EQ: Specifically, when the current state is IQ and the character passed at this time is a double quote, it is converted to EQ; the transition from IQ to IQ: Specifically, when the current state is IQ and the characters passed at this time are a comma, a line feed, and other characters, it is converted to IQ itself.
[0056] The transition from EQ to IQ: Specifically, when the current state is EQ and the character passed at this time is a double quote, it is converted to IQ; the transition from EQ to ER: Specifically, when the current state is EQ and the character passed at this time is a line feed, it is converted to ER; the transition from EQ to IR: Specifically, when the current state is EQ and the character passed at this time is a comma, it is converted to IR.
[0057] The record-level state transition and the field-level state transition are referred to as templates in the FST. The smallest set of possible state transitions in the CSV text is called the minimized state, and the states existing within the template are called the state set of the template in the FST.
[0058] Step 2.3: Finally, according to the control characters sampled in Step 1, the finite state transducer FST of the present invention is constructed from a finite set of states and corresponding transitions T. The state S is stored as an enumerated value, and multiple branch decisions are introduced for the state transition T.
[0059] A record is often considered as a complete line, that is Figure 1 the record level in Figure 1 ; when there is no double quotation mark sampled in Step 1, it is considered that the position of the delimiter can be determined through character conversion at the record level; when there is a double quotation mark sampled in Step 1, it is considered that it is necessary to combine
[0060] Further, Step 3 also consists of three parts: (preprocessing) the parallel mode for block scanning; (extraction) SIMD-accelerated bitwise extraction; (postprocessing) maintaining accurate indexes of records, fields, and symbols. Specifically, it includes the following steps:
[0061] Step 3.1: The present invention starts from the preprocessing stage, evenly divides the file, and initiates scanning of the file blocks. According to the paradigm of FST parallelism, this method allows the processing unit operations to perform pointer jumps to determine the state of each character in the block. Each processing unit starts from four states and interprets the transitions as final states.
[0062] The parsing of CSV text is often limited to instances where commas and line breaks are used as row and column delimiters. To achieve efficient record and field extraction, this method develops an instruction-level parallel solution for delimiter recognition using the vector formatting instructions of SIMD. Among them, a record refers to a line of CSV text, and a field refers to the CSV text within two comma delimiters.
[0063] Step 3.3: Maintaining accurate structural indexes: To facilitate the query progress, this method constructs table indexes (records and fields) through parallel bit operations and stores them in a compressed bitmap to form a bitmap index.
[0064] Further, the two query modes of Step 5 are specifically as follows:
[0065] Step 5.1: Keyword search query mode. The user provides the keyword to be queried, and through the method of the present invention, the row numbers and column numbers of all exactly matching items found in the file are returned;
[0066] Step 5.2: File combined query mode. The user provides the folder name to be queried and the keywords to be queried. All CSV files in the folder directory are read through the method of the present invention, and the file names, line numbers, and column numbers of all exactly matching items found in the files are returned.
[0067] Example:
[0068] The technical solution of the present invention is based on the state-minimized FST model to achieve parallelization and scalable processing of CSV files. On this basis, the present invention proposes two parallelization methods at the thread level and the instruction level, and the use of this method does not cause parsing ambiguity, and the state is designed based on the delimiter. Finally, a comparison of the parsing speed with the current common CSV file processing methods is carried out to verify the smooth execution of parallelization and its scalability under a large number of cores.
[0069] The present invention proposes a high-performance CSV file processing method based on the FST model. Compared with the common serial processing method of CSV files, the present invention uses the method of state conversion to compensate for the error handling caused by the inherent format of CSV files, realizes thread-level parallelism during the CSV file parsing process, that is, the user gives the available number of threads, automatically divides the blocks according to the number of threads, deploys SIMD vector formatting instructions inside each thread, and performs block scanning at the instruction level, parallelizing the CSV file from two levels: the thread level and the instruction level. At the same time, the present invention proposes an accurate structure index and an optional index optimization strategy. The specific technical solutions are as follows:
[0070] Step 1: Identify (sample) the control characters in the input file through an iterative method to determine the symbol state and logical position of the control characters;
[0071] Step 2: Determine the character conversion level that should be selected in the FST model according to the control characters;
[0072] Step 3: Divide the file into equal-sized text blocks and put them into idle processing units, perform parallel scanning based on FST, and use SIMD to accelerate the identification of delimiters to generate a bitmap index, mapping the logical positions of the delimiters to physical positions;
[0073] Step 4: Optionally, further process the bitmap index to simplify the query time, or directly perform the query without processing the bitmap index;
[0074] Step 5: Perform queries based on the constructed bitmap index, providing two query modes: keyword search and file combined query. Further, the present invention selects the custom_2020 dataset as an example, and the form of a certain record in its dataset is as follows:
[0075] 198801,1,103,100,000000190,0,35843,34353
[0076] The iterative identification symbol in Step 1 is specifically as follows:
[0077] Step 1.1: After the file is read in, according to the formula of the heuristic algorithm, the 101st to 132nd characters are sampled, a total of 32 characters, which conforms to the basic processing logic of SIMD;
[0078] Step 1.2: The 32 characters are compared with the control characters one by one. As recorded in this embodiment, when a comma separator is encountered, it is considered that the comma separator exists in the file, and the inspection of the comma separator is stopped. The detection of the remaining control characters is similar to that of the comma separator. In this embodiment, there are two types of separators, the comma and the line break. If only these two types of control characters are detected in the entire file, then the present invention considers that the control characters in the file are the comma and the line break, and then determines that the minimized states are IR and ER, and then mapping to the record-level state transition can meet the requirements.
[0079] Further, the FST character conversion level in Step 2 is specifically as follows:
[0080] Step 2.1: The present invention makes a one-to-one mapping between the state sets of different FST templates and the control characters, and its mapping relationship is that different control characters correspond to different state transitions, and vice versa does not hold;
[0081] Step 2.2: Obtain the set of minimum control characters from Step 1, match it with the templates in FST, and select the templates composed of the minimized states (i.e., IR and ER) and select them.
[0082] Further, the working process of Step 3 is specifically as follows:
[0083] Step 3.1: Read the entire file into memory, and according to the size of the file and the number of available threads, divide the file according to the number of available threads. For example, in this embodiment, the size of the file is 4.544 GB. Assuming there are 8 available threads, the block size is 568 MB, and for 16 available threads, the block size is 284 MB, and so on;
[0084] Step 3.2: Parallelly scan each block of the file. Inside the block, 32 characters are used as a group to adapt to the SIMD instruction set to achieve instruction-level parallelism, and state transitions are performed according to the selected FST template. For example, in this embodiment, the file only contains two types of control characters, the comma and the line break, which are mapped to the FST template, that is, the two states of IR and ER can cover all situations, and each block may finally only generate one of the four final states {IR, ER}, {IR, IR}, {ER, IR}, and {ER, ER};
[0085] Step 3.3: While parallelly scanning the chunks, construct bitmap indexes for the positions where commas and line breaks appear, namely the comma bitmap and the line break bitmap. The purpose is to map the positions of the delimiters onto the bitmaps, achieving the goal of reducing the occupied space and serving as the basis for further optimizing the index construction. It should be noted that at this time, the constructed bitmap indexes only represent the relative positions within the chunks, which can also be understood as logical positions rather than absolute positions, that is, physical positions;
[0086] Step 3.4: After all the chunks are processed, according to the final states given by each chunk, a single thread is used to serially connect the states between the chunks, and thus determine the physical positions of the delimiters. In this embodiment, bitmaps for determining the physical positions of the comma and line break delimiters respectively will be determined. And further processing is performed on this (i.e., Step 4).
[0087] Furthermore, the index processing in Step 4 is specifically as follows:
[0088] When further processing of the index is required, the present invention provides a method for further reducing the memory space occupied by the index and accelerating the query speed. In this embodiment, after the bitmap indexes for the comma and line break delimiters are constructed, the part of 0 between the indexes is considered as the value within the field, and at the same time, 0 is compressed, marking the delimiter positions and the field lengths;
[0089] Meanwhile, in the query stage, the present invention provides an optimization method based on memory occupancy. Align the file to continuous pages, and when querying to a position where the field length matches, read in the page where the field is located and perform character-by-character verification with the query value. This method has a greater impact in the query of long field lengths.
[0090] The key point of this application is: The present invention proposes a high-performance CSV processing method based on the FST model with minimized states. The present invention uses the method of state transition to compensate for the error handling caused by the inherent format of the CSV file; on this basis, the present invention proposes two parallelization methods at the thread level and the instruction level, and the use of this method does not cause parsing ambiguity, and the state design is based on control characters; finally, a comparison of the parsing speeds with the currently common CSV file processing methods is made to verify the smooth execution of the parallelization and the scalability under a large number of cores.
[0091] Another embodiment of the present invention provides a method and device for high-performance analysis of delimiter format files, including:
[0092] A control character sampling module, used to sample the control characters in the input file in an iterative manner to determine the symbol states and logical positions of the control characters;
[0093] A conversion level selection module, configured to determine the character conversion level to be selected in the finite state transducer model according to a control character, where the character conversion level includes a record level and a field level;
[0094] A bitmap index generation module, configured to divide an input file into text blocks of equal size, place them in idle processing units, implement parallel scanning based on the finite state transducer model, and utilize SIMD to accelerate the recognition of delimiters, so as to generate a bitmap index, where the bitmap index maps the logical positions of the delimiters to physical positions;
[0095] A query module, configured to perform queries based on the bitmap index, including a keyword search query mode and a file union query mode.
[0096] For the specific implementation processes of each module, refer to the description of the method of the present invention above.
[0097] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor, where the memory stores a computer program, and the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the method of the present invention.
[0098] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc), where the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps of the method of the present invention are implemented.
[0099] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A method for high-performance analysis of delimiter format files, characterized in that, It includes the following steps: Sample control characters in the input file in an iterative manner to determine the symbol status and logical position of the control characters; Determine the character conversion level to be selected in the finite state transducer model according to the control characters, and the character conversion level includes the record level and the field level; Slice the input file into text blocks of equal size and put them into idle processing units, implement parallel scanning based on the finite state transducer model, and use SIMD to accelerate the recognition of delimiters to generate a bitmap index, which maps the logical position of the delimiters to the physical position; Perform queries based on the bitmap index, including keyword search query mode and file union query mode; The iterative sampling of control characters in the input file includes using a heuristic algorithm to check the control characters in the file, that is, as long as a certain control character in the control characters is encountered, it is considered that the file contains control characters and the check of this control character is stopped until all characters in the control characters are included or the end of the file is reached; and, use the SIMD instruction set to quickly judge whether a continuous string in the CSV file contains a certain type of control character, while expanding the sampling range and maintaining the original sampling speed as much as possible; the formula of the heuristic algorithm is as follows: k < log2(length - i) where f(x) represents the position that the pointer should point to during each sampling, k represents the number of samplings, length represents the length of the text, i represents the position pointed to by the initial pointer, and the value of i is 100; The finite state transducer is a finite state transducer based on state minimization, including four states of IR, ER, IQ, and EQ for CSV format file conversion and their corresponding conversions, where IR represents within record, ER represents end of record, IQ represents within quote, and EQ represents end of quote; The conversions corresponding to the four states include: Conversion from IR to ER: When the current state is IR and the character passed at this time is a line feed character, it is converted to ER; Conversion from IR to IQ: When the current state is IR and the character passed at this time is a double quote, it is converted to IQ; Conversion from IR to IR: When the current state is IR and the character passed at this time is other characters except line feed characters and double quotes, it is converted to IR; Conversion from ER to IQ: When the current state is ER and the character passed at this time is a double quote, it is converted to IQ; Conversion from ER to IR: When the current state is ER and the symbol passed at this time is other characters, it is converted to IR; Conversion from IQ to EQ: When the current state is IQ and the character passed at this time is a double quote, it is converted to EQ; Conversion from IQ to IQ: When the current state is IQ and the character passed at this time is a comma, a line feed character, and other characters, it is converted to IQ itself; Conversion from EQ to IQ: When the current state is EQ and the character passed at this time is a double quote, it is converted to IQ; Conversion from EQ to ER: When the current state is EQ and the character passed at this time is a line feed character, it is converted to ER; Conversion from EQ to IR: When the current state is EQ and the character passed at this time is a comma, it is converted to IR; Determining the character conversion level to be selected in the finite state transducer model according to the control character includes: when there is no double quotation mark in the sampling, it is considered that the position of the delimiter can be determined through the character conversion at the record level; when there is a double quotation mark in the sampling, it is considered that the field level needs to be combined to perform character conversion at the field level to determine the position of the delimiter.
2. The method according to claim 1, characterized in that The control characters include delimiters, double quotation marks, and escape characters. The delimiters include commas and line breaks. The part other than the control characters is called text characters. The symbol state refers to whether it exists in the sampling. The logical position refers to the relative position of the sampled control character in the text.
3. The method according to claim 1, wherein Process the bitmap index to simplify the query time, or directly perform a query without processing the bitmap index.
4. The method according to claim 1, characterized in that, The keyword search query mode is that the user provides the keyword to be queried and returns the line numbers and column numbers of all exactly matching items found in the file. The file union query mode is that the user provides the folder name to be queried and the keyword to be queried, then reads all CSV files in the folder directory, and returns the file names, line numbers, and column numbers of all exactly matching items found in the file.
5. A method and apparatus for high-performance analysis of a delimiter format file using the method according to any one of claims 1 to 4, characterized in that, It includes: A control character sampling module for sampling the control characters in the input file in an iterative manner to determine the symbol state and logical position of the control characters. A conversion level selection module for determining the character conversion level to be selected in the finite state transducer model according to the control character. The character conversion level includes the record level and the field level. A bitmap index generation module for splitting the input file into text blocks of equal size, putting them into idle processing units, implementing parallel scanning based on the finite state transducer model, and using SIMD to accelerate the recognition of delimiters to generate a bitmap index. The bitmap index maps the logical position of the delimiter to the physical position. A query module for performing queries based on the bitmap index, including a keyword search query mode and a file union query mode.
6. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program is executed by the computer, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
General text format-oriented analysis method and tool
CN107341135A
Methods and apparatus to store and enable a transfer protocol for new file formats for distributed processing and preserve data integrity through distributed storage blocks
US10564877B1
Parser for Schema-Free Data Exchange Format
US20180314722A1