Method and system for generating a sparse suffix array
Generate sparse suffix arrays through preset value division and mapping sorting, which solves the problem of excessive space occupancy of suffix arrays and realizes efficient data retrieval and storage, especially suitable for large files and memory-constrained scenarios.
Patent Information
- Application Number
- CN202211520783.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-30
AI Technical Summary
Existing suffix arrays take up too much space in large files or memory-constrained situations, resulting in unsorted or generated suffix tree groups that cannot be imported into the computer for subsequent applications.
The source file is divided into multiple strings by preset values, and the threshold is configured according to the memory and file size. A sparse suffix array is generated using mapping and sorting algorithms, and the small-endian computer is arranged in reverse order to form an unsigned integer array, and finally a sparse suffix array is obtained.
Effectively saves running space, optimizes time and space complexity, is suitable for large files and memory-constrained environments, and improves data retrieval efficiency.
Smart Images

Figure CN116069989B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data structure construction, and particularly to a method and system for generating a sparse suffix array. Background Art
[0002] In computer science, a suffix array (SA) is an ordered array of all suffixes of a string (source file). It is a basic data structure for full-text indexing and has a wide range of applications in data retrieval, data compression, bibliometrics, and gene analysis. A typical use case of SA is to find the longest substring match in the source file for a query string.
[0003] For the existing suffix array SA, each position in the source file needs log2n bits (n is the number of bytes of the source file), so the space required for the suffix array SA is bytes; it can be seen that the suffix array SA occupies too much space, and it may be difficult to load the entire suffix array SA into the computer memory for a large file. In other words, when the source file is too large or the computing memory is limited, it is impossible to sort all the suffixes, or even if all the suffixes can be sorted, the generated suffix array is very likely not to be imported into the computer for subsequent applications due to space limitations. Summary of the Invention
[0004] In view of the problems existing in the prior art, the present invention provides a method for generating a sparse suffix array, which pre-configures a preset value according to the memory of a computing terminal and the file size of a source file;
[0005] The generation method includes:
[0006] Step S1, the computing terminal divides the source file into multiple strings according to the preset value, and the length of each string is the preset value;
[0007] Step S2, the computing terminal determines whether the preset value is greater than a threshold:
[0008] If so, sort each string to obtain a corresponding string sorting result, and then turn to step S3;
[0009] If not, turn to step S4;
[0010] Step S3, the computing terminal maps each corresponding string in the string sorting result to an integer according to a pre-configured mapping relationship to form an unsigned integer array, and then turns to step S5;
[0011] Step S4, the computing terminal obtains its own attribute parameters, and determines whether it is a little-endian computer according to the attribute parameters:
[0012] If not, convert each of the strings into an integer respectively to form an unsigned integer array, and then proceed to step S5;
[0013] If so, reverse the order of each of the strings to obtain reverse strings, and convert each of the reverse strings into an integer to form an unsigned integer array, and then proceed to step S5;
[0014] Step S5, the computing terminal performs integer suffix sorting on the unsigned integer array to obtain a sorted array, and corrects each element in the sorted array according to the preset value to obtain the sparse suffix array of the source file.
[0015] Preferably, in step S1, it includes:
[0016] Step S11, the computing terminal sequentially selects characters of the preset value length from the source file to form each of the strings, and counts the string length of the last string;
[0017] Step S12, the computing terminal determines whether the string length reaches the preset value:
[0018] If so, proceed to step S2;
[0019] If not, correct the last string, and then proceed to step S2.
[0020] Preferably, in step S12, correcting the last string includes deleting the last string, or performing zero-padding after the last string until the string length of the last string reaches the preset value.
[0021] Preferably, the threshold is 4 or 8.
[0022] Preferably, in step S3, the mapping relationship is that the sequence order of the string sorting results is respectively mapped to each of the integers in an arithmetic sequence with 0 as the first term and 1 as the common difference.
[0023] Preferably, in step S5, correcting each element in the sorted array according to the preset value is to multiply each element in the sorted array by the preset value.
[0024] The present invention provides a sparse suffix array generation system, which applies the above generation method. The generation system is installed in a computing terminal, and the generation system includes:
[0025] A file division module, configured to divide a source file into multiple strings according to a preset value, and the length of each of the strings is the preset value;
[0026] A first judgment module, connected to the file division module, configured to, when judging that the preset value is greater than a threshold, sort each of the strings to obtain a corresponding string sorting result, and generate an attribute judgment signal when the preset value is not greater than the preset value;
[0027] A first sorting module, connected to the first judgment module, configured to respectively map each of the corresponding strings in the string sorting result to an integer according to a pre-configured mapping relationship to form an unsigned integer array;
[0028] A second judgment module, connected to the first judgment module, configured to obtain attribute parameters of the carried computing terminal according to the attribute judgment signal, and when judging that the computing terminal is a big-endian computer according to the attribute parameters, respectively convert each of the strings into an integer to form an unsigned integer array, and when judging that the computing terminal is a little-endian computer, respectively reverse each of the strings to obtain reversed strings, and convert each of the reversed strings into an integer to form an unsigned integer array;
[0029] A second sorting module, respectively connected to the first sorting module and the second judgment module, configured to perform integer suffix sorting on the unsigned integer array to obtain a sorted array, and correct each element in the sorted array according to the preset value to obtain a sparse suffix array of the source file.
[0030] Preferably, the threshold is 4 or 8.
[0031] Preferably, the mapping relationship is that the order of the string sorting result is respectively mapped to each of the integers in an arithmetic progression with 0 as the first term and 1 as the common difference.
[0032] Preferably, the second sorting module corrects each element in the sorted array according to the preset value by multiplying each element in the sorted array by the preset value.
[0033] The above technical solution has the following advantages or beneficial effects: only sort the suffix array of the source file starting from each preset value position, and the space required for the sparse suffix array of the present technical solution is (k represents the preset value, and n represents the number of bytes of the source file), saving at least the preset value times the space compared with the existing suffix array. Theoretically, the optimal time complexity of sparse suffix sorting is O(n), and the space complexity is O(n / k). The present technical solution reaches the optimal in both algorithm time and space complexity. Description of the Drawings
[0034] Figure 1 In a preferred embodiment of the present invention, it is a schematic flowchart of a method for generating a sparse suffix array;
[0035] Figure 2 In a preferred embodiment of the present invention, it is a schematic sub - flowchart of step S1;
[0036] Figure 3 In a preferred embodiment of the present invention, it is a schematic structural diagram of a system for generating a sparse suffix array. Detailed implementation manners
[0037] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The present invention is not limited to this embodiment, and as long as it conforms to the gist of the present invention, other embodiments also belong to the scope of the present invention.
[0038] In a preferred embodiment of the present invention, in view of the above - mentioned problems existing in the prior art, a method for generating a sparse suffix array is provided. A preset value is configured in advance according to the memory of a computing terminal and the file size of a source file;
[0039] As Figure 1 shown, the generation method includes:
[0040] Step S1, the computing terminal divides the source file into multiple strings according to the preset value, and the length of each string is the preset value;
[0041] Step S2, the computing terminal determines whether the preset value is greater than a threshold:
[0042] If so, sort each string to obtain a corresponding string sorting result, and then turn to step S3;
[0043] If not, turn to step S4;
[0044] Step S3, the computing terminal maps each corresponding string in the string sorting result to an integer according to a pre - configured mapping relationship to form an unsigned integer array, and then turns to step S5;
[0045] Step S4, the computing terminal obtains its own attribute parameters and determines whether it is a little - endian computer according to the attribute parameters:
[0046] If not, convert each of the said strings into an integer to form an unsigned integer array, and then turn to step S5;
[0047] If so, reverse - order each string to obtain a reverse - ordered string, and convert each of the said reverse - ordered strings into an integer to form an unsigned integer array, and then turn to step S5;
[0048] Step S5, calculate the sorted array obtained by the integer suffix sorting of the unsigned integer array by the terminal, and correct each element in the sorted array according to a preset value to obtain the sparse suffix array of the source file.
[0049] Specifically, taking the preset value as k, in this embodiment, only the suffix array of the source file starting from every k positions is sorted. The space required for the sparse suffix array SSA of this technical solution is (k represents the preset value, and n represents the number of bytes of the source file), which saves at least k times the space compared to the implementation of the existing suffix array sorting. Among them, the size of k can be determined according to the actual situation, including but not limited to determining the size of k according to the available memory space of the computing terminal or the accuracy of suffix sorting, etc. Preferably, when the number of bytes n of the source file and the available memory space M of the computing terminal are known, under the premise of satisfying , the smaller the value of k, the better. Theoretically, the optimal time complexity of sparse suffix sorting is O(n), and the space complexity is O(n / k). This technical solution achieves the optimal in both algorithm time and space complexity.
[0050] Further specifically, after the source file is divided into strings, according to different configurations of the preset value, the lengths of each string are different. If the preset value is not greater than the threshold, it means the string is short, and each string can be regarded as an unsigned long integer without further processing. When the preset value is greater than the threshold, it means the string is long and cannot be regarded as an unsigned long integer, and then it needs to be sorted and mapped in sequence so that each sorted string corresponds to an integer respectively, and finally an unsigned integer array is obtained. Among them, for the case where the preset value is not greater than the threshold, considering that a big-endian computer stores the high bits at the low address when storing data, while a little-endian computer stores the high bits at the high address when storing data. Therefore, for a little-endian computer, it is necessary to reverse the order of each string respectively to determine the value of the unsigned long integer actually represented by each corresponding string.
[0051] Among them, in step S2, when the preset value is greater than the threshold, it is preferably to use the Most Significant Digit First Radix Sort (MSDRadixSort) to sort each string to obtain the corresponding string sorting result. The Most Significant Digit First Radix Sort not only has been among the top in terms of speed and uses only 9n / k of space (k represents the preset value, and n represents the number of bytes of the source file); but also performs very well for large files and multi-core CPUs. In step S5, an existing integer suffix sorting algorithm can be used to implement the processing of the unsigned integer data to the sorted array, and the implementation process of the specific sorting algorithm is not the inventive point of this technical solution and will not be elaborated here.
[0052] Further preferably, the technical solution can be applied to the differential algorithm, which is the core algorithm of incremental update. It describes the calculation of the differences between the new and old version files, thus greatly reducing the size of the update package. It needs to start matching the old file from any position of the new file. Based on this, the method for generating the sparse suffix array of the present technical solution can be called in the differential algorithm to generate the sparse suffix array of the old file, improving the matching efficiency between the new file and the old file while having a low requirement for space, especially for old files with a relatively large amount of data, and having strong applicability.
[0053] Furthermore, a large number of experiments on sparse suffix arrays were conducted in the differential algorithm, and it was found that for the typical values k = 4 or 8, the average increase in the update package size is ~2% and ~4%. That is to say, the impact of the decrease in the query effect of using the sparse suffix of the present technical solution compared to using the full-text suffix is limited.
[0054] In a preferred embodiment of the present invention, as Figure 2 shown, step S1 includes:
[0055] Step S11, calculating the terminal to sequentially select characters with a preset length from the source file to form each string, and counting the string length of the last string;
[0056] Step S12, the terminal judges whether the string length reaches the preset value:
[0057] If so, go to step S2;
[0058] If not, correct the last string, and then go to step S2.
[0059] In a preferred embodiment of the present invention, in step S12, correcting the last string includes deleting the last string or padding with 0s after the last string until the string length of the last string reaches the preset value.
[0060] In a preferred embodiment of the present invention, the threshold is 4 or 8.
[0061] In a preferred embodiment of the present invention, in step S3, the mapping relationship is that the sequence order of the string sorting results is sequentially mapped to each integer in the arithmetic sequence with 0 as the first term and 1 as the common difference.
[0062] In a preferred embodiment of the present invention, in step S5, correcting each element in the sorted array according to the preset value is to multiply each element in the sorted array by the preset value.
[0063] The present invention provides a system for generating a sparse suffix array, which applies the above-mentioned generation method. The generation system is installed in a computing terminal, as Figure 3As shown, the generation system includes:
[0064] A file division module 1, configured to divide a source file into multiple strings according to a preset value, and the length of each string is the preset value;
[0065] A first judgment module 2, connected to the file division module 1, configured to sort each string to obtain a corresponding string sorting result when it is judged that the preset value is greater than a threshold value, and generate an attribute judgment signal when the preset value is not greater than the preset value;
[0066] A first sorting module 3, connected to the first judgment module 2, configured to respectively map each corresponding string in the string sorting result to an integer according to a pre-configured mapping relationship to form an unsigned integer array;
[0067] A second judgment module 4, connected to the first judgment module 2, configured to obtain the attribute parameters of the carried computing terminal according to the attribute judgment signal, and when it is judged that the computing terminal is a big-end computer according to the attribute parameters, respectively convert each string into an integer to form an unsigned integer array, and when it is judged that the computing terminal is a little-end computer, respectively reverse each string to obtain a reversed string, and convert each reversed string into an integer to form an unsigned integer array;
[0068] A second sorting module 5, respectively connected to the first sorting module 3 and the second judgment module 4, configured to perform integer suffix sorting on the unsigned integer array to obtain a sorted array, and correct each element in the sorted array according to the preset value to obtain the sparse suffix array of the source file.
[0069] In a preferred embodiment of the present invention, the threshold value is 4 or 8.
[0070] In a preferred embodiment of the present invention, the mapping relationship is that the sequence of the string sorting result is respectively mapped to each integer in an arithmetic sequence with 0 as the first term and 1 as the common difference.
[0071] In a preferred embodiment of the present invention, the second sorting module 5 corrects each element in the sorted array according to the preset value by multiplying each element in the sorted array by the preset value.
[0072] The above are only preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be able to realize that all equivalent replacements and obvious changes made by using the content of this specification and the drawings should be included in the protection scope of the present invention.
Claims
1. A method for generating a sparse suffix array, characterized in that, Pre-configure a preset value according to the memory of a computing terminal and the file size of a source file; Then the generating method includes: Step S1, the computing terminal divides the source file into multiple strings according to the preset value, and the length of each string is the preset value; Step S2, the computing terminal determines whether the preset value is greater than a threshold: If so, sort each of the strings to obtain a corresponding string sorting result, and then proceed to step S3; If not, proceed to step S4; Step S3, the computing terminal maps each of the corresponding strings in the string sorting result to an integer respectively according to a pre-configured mapping relationship to form an unsigned integer array, and then proceeds to step S5; Step S4, the computing terminal obtains its own attribute parameters, and determines whether it is a little-endian computer according to the attribute parameters: If not, convert each of the strings into an integer respectively to form an unsigned integer array, and then proceed to step S5; If so, reverse each of the strings to obtain a reversed string, and convert each of the reversed strings into an integer to form an unsigned integer array, and then proceed to step S5; Step S5, the computing terminal performs integer suffix sorting on the unsigned integer array to obtain a sorted array, and corrects each element in the sorted array according to the preset value to obtain the sparse suffix array of the source file.
2. The generation method according to claim 1, characterized in that In the step S1, it includes: Step S11, the computing terminal sequentially selects characters with the length of the preset value from the source file to form each of the strings, and counts the string length of the last string; Step S12, the computing terminal determines whether the string length reaches the preset value: If so, proceed to the step S2; If not, correct the last string, and then proceed to the step S2.
3. The generation method according to claim 2, wherein In the step S12, correcting the last string includes deleting the last string, or performing zero-padding after the last string until the string length of the last string reaches the preset value.
4. The generation method according to claim 1, characterized in that, The threshold is 4 or 8.
5. The generation method according to claim 1, wherein In the step S3, the mapping relationship is that the sequence order of the string sorting result is correspondingly mapped to each of the integers in an arithmetic sequence with 0 as the first term and 1 as the common difference.
6. The generation method according to claim 1, wherein In the step S5, correcting each element in the sorted array according to the preset value is to multiply each element in the sorted array by the preset value.
7. A sparse suffix array generation system, characterized in that Applying the generating method according to any one of claims 1-6, the generating system is installed in the computing terminal, and the generating system includes: A file division module, configured to divide a source file into multiple strings according to a preset value, and the length of each string is the preset value; A first judgment module, connected to the file division module, configured to sort each of the strings to obtain a corresponding string sorting result when it is judged that the preset value is not greater than a threshold, and generate an attribute judgment signal when the preset value is greater than the preset value; The first sorting module, connected to the first judgment module, is configured to map each of the corresponding strings in the string sorting result to an integer according to a pre-configured mapping relationship, so as to obtain an unsigned integer array; The second judgment module, connected to the first judgment module, is configured to obtain the attribute parameters of the carried computing terminal according to the attribute judgment signal, and when it is judged that the computing terminal is a little-endian computer according to the attribute parameters, reverse each of the strings to obtain reversed strings, and convert each of the reversed strings into an integer to form an unsigned integer array; The second sorting module is respectively connected to the first sorting module and the second judgment module, and is configured to perform integer suffix sorting on the unsigned integer array to obtain a sorted array, and correct each element in the sorted array according to the preset value to obtain the sparse suffix array of the source file.
8. The generation system according to claim 7, wherein The threshold is 4 or 8.
9. The generating system according to claim 7, wherein The mapping relationship is that the order of the string sorting result is sequentially mapped to each of the integers in an arithmetic sequence with 0 as the first term and 1 as the common difference.
10. The generation system according to claim 7, wherein The second sorting module corrects each element in the sorted array according to the preset value by multiplying each element in the sorted array by the preset value.
Citation Information
Patent Citations
Method of self-adaptive merging of suffix arrays and device thereof
CN108664459A
Method and system for constructing suffix arrays (SAs) in parallel in constant working space
CN108763170A