Semi-structured data compression and storage method, device and equipment and storage medium

By adopting columnar storage and specific compression methods in semi-structured data processing, the problems of inconvenience in reading and under-optimal compression performance in the prior art are solved, efficient data compression and fast data reading are achieved, and data readability and storage efficiency are improved.

CN119988680APending Publication Date: 2025-05-13BEIJING YOUTEJIE INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510109819.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When processing and storing massive semi-structured data in the prior art, there are problems such as inconvenience in reading and failure to achieve optimal compression performance.

Method used

Using columnar storage and specific compression methods, a mapping table is generated by parsing semi-structured data, the data structure of columnar storage files is determined, and the data is compressed and stored in columnar storage files based on the header information.

Benefits of technology

It realizes that the values ​​of specific fields can be read without decompressing the entire piece of data, improves data reading efficiency, improves data readability and convenience of use, and achieves a better compression effect than traditional general compression algorithms, saving storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988680A_ABST
    Figure CN119988680A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-structured data compression storage method and device, equipment and a storage medium. Comprising the following steps: acquiring original semi-structured data, and analyzing the original semi-structured data to generate a mapping table; determining a data structure of the column storage file according to the mapping table, and creating header information of the column storage file according to the data structure; and compressing and storing the original semi-structured data into the column storage file based on the header information. By adopting the column storage and the specific compression mode, the value of the specific field can be read without decompressing the whole segment of data, so that the data reading efficiency is greatly improved. The reading inconvenience caused by integral compression is avoided, the data is more flexible and convenient to use, and the readability of the data is improved. The characteristics of the semi-structured data are fully utilized for compression, the compression effect better than that of a traditional general compression algorithm can be achieved, the storage space is saved, and the hardware and time cost of enterprises in the aspects of data storage and processing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a semi-structured data compression storage method, device, equipment and storage medium. Background Art

[0002] In today's era of explosive data growth, semi-structured data represented by JSON is widely used in many fields such as log data, user behavior data, sensor data, etc., because it does not require pre-defined field names and data types and has self-describing characteristics. Its importance is self-evident. However, as the scale of data continues to expand, enterprises face tremendous pressure in processing and storing massive semi-structured data.

[0003] At present, a general compression algorithm is commonly used for the storage of semi-structured data, that is, by analyzing the repeated patterns in the data to generate a dictionary table, and then replacing the repeated parts of the original data with dictionary indexes, and reducing the data volume through encoding and compression. Although it has certain performance in terms of storage efficiency, since the data is compressed as a whole, the entire segment of data must be decompressed before reading a certain field value, which has obvious limitations in readability and ease of use. In addition, the existing technology cannot fully utilize the characteristics of semi-structured data that are already partially structured, so that the compression performance fails to reach the optimal state. Summary of the invention

[0004] The present invention provides a semi-structured data compression storage method, device, equipment and storage medium, which can achieve efficient data compression while maintaining data readability and ease of use.

[0005] According to one aspect of the present invention, a method for compressing and storing semi-structured data is provided, the method comprising:

[0006] Acquire original semi-structured data, and parse the original semi-structured data to generate a mapping table;

[0007] Determine the data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure;

[0008] The original semi-structured data is compressed and stored in columnar storage files based on the header information.

[0009] Optionally, the original semi-structured data is parsed to generate a mapping table, including: traversing each field in the original semi-structured data; performing key-value mapping on each field to construct an element correspondence relationship, and generating a mapping table according to the element correspondence relationship.

[0010] Optionally, determining the data structure of the column storage file according to the mapping table includes: determining the data type of each element in the mapping table, where the data type includes a string, an integer, a list, and a nested table; determining the format specification of the corresponding column storage file according to the data type, and constructing a Schema field definition based on the format specification; integrating each Schema field definition to generate a data structure.

[0011] Optionally, the original semi-structured data is compressed and stored in a column storage file based on the header information, including: traversing each field in the original semi-structured data and its corresponding field data; taking each field as a target field in turn, and determining the target header information corresponding to the target field in the header information; determining the target column data block corresponding to the target header information from the column storage file; and using a specified compression algorithm to store the target field data corresponding to the target field in the target column data block.

[0012] Optionally, after compressing and storing the original semi-structured data in a columnar storage file based on the header information, the method further includes: determining attribute information of each column data block in the columnar storage file, wherein the attribute information includes a column value range and a row group; and establishing index information for each column data block based on the attribute information.

[0013] Optionally, the method also includes: determining the target data access mode and the target storage device type; obtaining a storage policy list, wherein the storage policy list includes storage policies corresponding to each data access mode and storage device type; matching the target data access mode and the target storage device type through the storage policy list to obtain a corresponding target storage policy, wherein the target storage policy includes data block capacity and compression algorithm type.

[0014] Optionally, the method further includes: performing storage optimization on the columnar storage files according to a specified period, wherein the storage optimization includes merging files, reorganizing data blocks, and updating index information.

[0015] According to another aspect of the present invention, there is provided a semi-structured data compression storage device, the device comprising:

[0016] A data parsing and mapping module is used to obtain original semi-structured data and parse the original semi-structured data to generate a mapping table;

[0017] A header information generation module is used to determine the data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure;

[0018] The data compression storage module is used to compress and store the original semi-structured data into a column storage file based on the header information.

[0019] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0020] at least one processor;

[0021] and a memory communicatively coupled to the at least one processor;

[0022] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a semi-structured data compression storage method described in any embodiment of the present invention.

[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a semi-structured data compression storage method described in any embodiment of the present invention when executed.

[0024] The technical solution of the embodiment of the present invention, by adopting column storage and a specific compression method, can read the value of a specific field without decompressing the entire data, thereby greatly improving the efficiency of data reading. It avoids the inconvenience of reading caused by overall compression, makes the use of data more flexible and convenient, and improves the readability of data. By making full use of the characteristics of semi-structured data for compression, it can achieve a better compression effect than traditional general compression algorithms, save storage space, and reduce the hardware and time costs of enterprises in data storage and processing.

[0025] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 is a flow chart of a semi-structured data compression storage method provided according to the first embodiment of the present invention;

[0028] Figure 2 is a flowchart of another semi-structured data compression storage method provided according to the second embodiment of the present invention;

[0029] Figure 3is a structural schematic diagram of a semi-structured data compression storage device provided according to Embodiment 3 of the present invention;

[0030] Figure 4 It is a structural schematic diagram of an electronic device for implementing a semi-structured data compression storage method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] Embodiment 1

[0034] Figure 1 A flowchart of a semi-structured data compression storage method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of semi-structured data storage. The method can be executed by a semi-structured data compression storage device. The semi-structured data compression storage device can be implemented in the form of hardware and / or software. The semi-structured data compression storage device can be configured in a computer controller. Figure 1 As shown, the method includes:

[0035] S110: Acquire original semi-structured data, and parse the original semi-structured data to generate a mapping table.

[0036] Among them, semi-structured data refers to data between structured data and unstructured data. It has a certain structure, but it is not as regular as structured data, such as JSON and XML. It has tags or keys that identify the meaning of the data, but the arrangement and type definition of the data are relatively flexible. Parsing refers to analyzing the original semi-structured data according to its specific format rules and breaking it into understandable components or elements for further processing. A mapping table refers to a data structure, which can be a Map structure in the Java language, which records the correspondence between each element in the original semi-structured data and the storage location, storage method or other related information, and can provide guidance for subsequent storage. For example, when parsing JSON data, the mapping table records the data type, hierarchy and other information of the value corresponding to each key. The hierarchical relationship in the mapping table can help determine which data elements belong to the same logical column.

[0037] Specifically, the parsing process depends on the grammatical rules of the data format. For example, for JSON data, the parser will recognize that curly braces represent objects, square brackets represent arrays, colons are used to separate key-value pairs, etc. During the parsing process, the parser traverses the data structure. Assuming the JSON data is {"name": "John", "age": 30, "hobbies": ["reading", "swimming"]}, the parser will record that the value corresponding to "name" is the string "John", the value corresponding to "age" is the number "30", and "hobbies" corresponds to an array whose elements are the strings "reading" and "swimming".

[0038] S120: Determine a data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure.

[0039] Among them, column storage files refer to files that store data by column. By storing data in the same column together, it is convenient for specific queries and analysis, such as data compression and aggregate calculations. For example, for a table containing names, ages, and addresses, column storage will store all names, all ages, and all addresses separately. The data structure refers to the organization of data in a column storage file, which is determined by a mapping table and includes information such as the data type of each column and the relationship between columns, which determines how the data is stored and accessed. The header information refers to the metadata located at the header of the column storage file that describes the characteristics of each column in the file, including the column name, data type, column attributes, etc.

[0040] Optionally, determining the data structure of the column storage file according to the mapping table includes: determining the data type of each element in the mapping table, where the data type includes a string, an integer, a list, and a nested table; determining the format specification of the corresponding column storage file according to the data type, and constructing a Schema field definition based on the format specification; integrating each Schema field definition to generate a data structure.

[0041] It should be noted that after the JSON data is parsed, the data type mapping and conversion module starts working. Since the data type of JSON data is relatively flexible, and the Parquet format has a strict definition of the data type, it is necessary to map the type of JSON data to a Parquet-compatible data type. The core operation of this link is to conduct a comprehensive and detailed traversal of the parsed mapping table, deeply analyze the data type of each element, and build a schema that accurately matches it according to the strict specifications of the Parquet format. Parquet Schema, as the data organization structure definition of the entire Parquet file, specifies in detail the core attributes of each column of data, such as the name, data type, whether null values ​​are allowed, and the nested structure of complex data.

[0042] Specifically, the data type identification can be determined based on the field identifier of the element. When the "type" field of the element is clearly identified as "string", such as the "name" field, its data type can be determined to be a string. The string type is usually used to store text information, and the string encoding rules (such as UTF-8) are followed during storage and processing. When the "type" field of the element is clearly identified as "number" and the value is in integer form, such as the "age" field, it is identified as an integer type. Integers are represented in a specific binary format when stored, and different programming languages ​​and storage systems may have different integer representation ranges and precisions. For fields such as "hobbies", when the "type" field is "array", it is determined to be a list type. List types can contain multiple elements of the same type. When the "type" field of an element is "object" and contains a "children" field, it indicates that the element is a nested structure. For example, the "address" field has child elements such as "city" and "country" inside it, forming a nested structure.

[0043] In a specific implementation, when building the Parquet Schema, each element in the jsonMap is traversed, the instanceof keyword is used to obtain the type of the value, and the corresponding Parquet Schema field definitions are built for different types (such as strings, integers, lists, nested maps, etc.). For list types, its element type is further checked to determine the appropriate Parquet field type and marked as repeated; for nested map types, the inner Schema is recursively built and added to the outer Schema.

[0044] S130: compress the original semi-structured data based on the header information and store it in a column storage file.

[0045] Among them, compressed storage refers to a storage method that uses compression algorithms to reduce the storage space occupied by data. By removing redundant information in the data and using more compact encoding to represent the data, such as Huffman encoding and LZ77 algorithm, it saves space and bandwidth when storing and transmitting data.

[0046] Specifically, the header information can provide guidance for data storage, that is, the controller can store the corresponding values ​​in the original semi-structured data into the corresponding columns according to the column order and data type defined in the header. For example, all name values ​​are stored in the storage location corresponding to the "name" column, and the age value is stored in the location corresponding to the "age" column. If the data has a nested or complex structure, it can be expanded and reasonably stored in the corresponding column according to the mapping table and the determined data structure.

[0047] In a specific implementation, Huffman coding can construct a coding table based on the frequency distribution of data, using shorter codes for data with high frequency and longer codes for data with low frequency. When storing data, first count the frequency of occurrence of different values ​​in each column of data, then generate a coding table based on the Huffman algorithm, and replace the data value with the corresponding code for storage. In addition, the compression algorithm can also operate in reverse when reading data, and restore the original data according to the coding table. During actual storage, the compressed data will be stored in a column storage file together with the header information in a certain format. For example, the header information can be stored first, and then the compressed data of each column can be stored in sequence.

[0048] Optionally, the original semi-structured data is compressed and stored in a column storage file based on the header information, including: traversing each field in the original semi-structured data and its corresponding field data; taking each field as a target field in turn, and determining the target header information corresponding to the target field in the header information; determining the target column data block corresponding to the target header information from the column storage file; and using a specified compression algorithm to store the target field data corresponding to the target field in the target column data block.

[0049] Specifically, the column storage file organizes data by column, and the data of each column is stored in a specific data block. The field name and other information in the target header information can be used to locate the corresponding column data block in the column storage file. For example, for the target header information corresponding to the name field, it may contain information such as the offset and length of the name column data block in the file. Through the target header information, the target column data block corresponding to the name field can be accurately located from the column storage file.

[0050] In a specific implementation, the controller can traverse all JSON data according to the column storage method of Parquet, and write the data of each field into the corresponding column data block in sequence. During the writing process, the column data is compressed using an efficient compression algorithm supported by Parquet, such as Lz4 or Zstd. For example, for a field column containing a large number of repeated strings, the compression algorithm can effectively remove duplicate data, thereby greatly reducing storage space.

[0051] Optionally, after compressing and storing the original semi-structured data in a columnar storage file based on the header information, the method further includes: determining attribute information of each column data block in the columnar storage file, wherein the attribute information includes a column value range and a row group; and establishing index information for each column data block based on the attribute information.

[0052] It should be noted that in order to facilitate data reading and querying, index information is also established in the column storage file, such as an index based on column value range or an index based on row group, so that the required data block can be quickly located when querying data.

[0053] Specifically, the value range can be determined by traversing all values ​​stored in the column data block. For example, for a column containing integer type data, a maximum value and a minimum value are initially set as the initial range. Traverse each integer in the data block. If a number smaller than the current minimum value is encountered, the minimum value is updated; if a number larger than the current maximum value is encountered, the maximum value is updated. A row group is a logical unit that divides data by row in order to optimize query and storage management. Determining a row group can be based on a fixed number of rows or on the amount of data. For example, set every 1,000 rows as a row group. When data is stored by column, starting from the first row, each 1,000 rows of data are divided into a row group. When storing, information such as the starting position and end position of each row group in the column data block is recorded.

[0054] For example, when the column value range is 78-95, the column value range can be divided into multiple sub-ranges, such as 70-80, 81-90, and 91-100. Then, the row group or the position of the column data block where the data in each sub-range is located is recorded. When querying data in a certain range, the index can be used to quickly locate the location that may contain the target data, reducing the need to scan the entire table.

[0055] Optionally, the method also includes: determining the target data access mode and the target storage device type; obtaining a storage policy list, wherein the storage policy list includes storage policies corresponding to each data access mode and storage device type; matching the target data access mode and the target storage device type through the storage policy list to obtain a corresponding target storage policy, wherein the target storage policy includes data block capacity and compression algorithm type.

[0056] Among them, the data access mode describes the characteristics such as how and how often data is read or written. Common data access modes include random access and sequential access. Storage device types include hard disk drives (HDDs), solid-state drives (SSDs), and flash memory. For example, HDDs have relatively low costs and large capacities, but slow read and write speeds, making them suitable for cost-sensitive scenarios that do not require high read and write speeds; SSDs have fast read and write speeds, but are more expensive and have relatively small capacities, making them suitable for scenarios that require high read and write performance. The controller can determine the target storage device type based on the application's requirements for storage performance, cost, capacity, and other factors.

[0057] Specifically, the storage policy list includes storage policies corresponding to each data access mode and storage device type. The storage policy contains data storage methods optimized for specific data access modes and storage device types, including data block capacity settings, compression algorithm selection, etc. For example, the list may stipulate: for random access mode and the storage device is SSD, a smaller data block capacity is used to improve random read and write performance, and a compression algorithm with a moderate compression ratio and fast decompression speed is selected; for sequential access mode and the storage device is HDD, a larger data block capacity is used to improve sequential read and write efficiency, and a compression algorithm with a high compression ratio is selected. The controller can use the determined target data access mode and target storage device type as query conditions to search for matching records in the storage policy list.

[0058] It should be noted that the target storage strategy includes data block capacity and compression algorithm type. The choice of data block capacity affects storage and access efficiency. Smaller data blocks are suitable for random access, which can reduce the amount of data read and written each time and improve response speed; larger data blocks are suitable for sequential access, which can reduce the number of I / O operations and improve overall read and write efficiency. The choice of compression algorithm type requires comprehensive consideration of factors such as compression ratio, compression and decompression speed. For example, for data that needs to be decompressed quickly, a compression algorithm with a fast decompression speed may be selected; for scenarios with limited storage capacity, an algorithm with a high compression ratio may be preferred.

[0059] Optionally, the method further includes: performing storage optimization on the columnar storage files according to a specified period, wherein the storage optimization includes merging files, reorganizing data blocks, and updating index information.

[0060] It should be noted that during the process of continuous data writing and operation, column storage files may generate multiple small files. Small files increase the management overhead of the storage system, and when reading data, files may need to be switched frequently, reducing I / O efficiency. Merging files is to merge small files into larger files to improve storage and access efficiency.

[0061] In addition, with the insertion, deletion and update of data, the data distribution within the data block may become no longer compact or orderly, affecting the access efficiency of the data. Reorganizing the data block can optimize the storage layout of the data within the block to improve the data reading performance.

[0062] Furthermore, since merging files and reorganizing data blocks changes the storage location and structure of data, the original index information is no longer accurate, so the index needs to be updated to ensure that data can be correctly and quickly located. By performing storage optimization operations at a specified period, the performance of columnar storage files can be effectively improved, ensuring that data is always stored and accessed efficiently during long-term use.

[0063] The technical solution of the embodiment of the present invention, by adopting column storage and a specific compression method, can read the value of a specific field without decompressing the entire data, thereby greatly improving the efficiency of data reading. It avoids the inconvenience of reading caused by overall compression, makes the use of data more flexible and convenient, and improves the readability of data. By making full use of the characteristics of semi-structured data for compression, it can achieve a better compression effect than traditional general compression algorithms, save storage space, and reduce the hardware and time costs of enterprises in data storage and processing.

[0064] Embodiment 2

[0065] Figure 2 This is a flowchart of a semi-structured data compression storage method provided in the second embodiment of the present invention. This embodiment adds a specific process of parsing the original semi-structured data to generate a mapping table based on the above-mentioned first embodiment. Among them, the specific content of steps S240-S250 is roughly the same as steps S120-S130 in the first embodiment, so they will not be repeated in this embodiment. Figure 2 As shown, the method includes:

[0066] S210: Obtain original semi-structured data.

[0067] S220 , traverse each field in the original semi-structured data.

[0068] S230 , performing key-value mapping on each field to construct an element correspondence relationship, and generating a mapping table according to the element correspondence relationship.

[0069] In a specific implementation, the open source Jackson library can be used to parse JSON data into a Java language Map structure. For example, for a JSON data containing user information, the module introduces the ObjectMapper class in the Jackson library. This core class undertakes the key task of JSON data parsing. By calling the readValue method of ObjectMapper and passing in jsonData and TypeReference <Map<String,Object> >Type reference, you can parse JSON data into Map<String,Object> Structure. During the parsing process, Jackson will automatically map the key-value pairs in JSON to the corresponding elements of the Java Map according to the syntax rules and structural characteristics of the JSON data. For example, in the above JSON data, the value "John" corresponding to the "name" key will be stored under the "name" key in the Map, and its value is the string type "John"; the value array [85,60,90] corresponding to the "bwh" key will be stored as a list object under the "bwh" key in the Map, and the elements in the list are integer values; and the value corresponding to the "other" key is a nested JSON object {"sex": "male", "age": 30}, which will become a sub-Map under the "other" key in the Map after parsing, containing the two sub-keys "sex" and "age" and the corresponding string "male" and integer value 30.

[0070] S240: Determine the data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure.

[0071] Optionally, determining the data structure of the column storage file according to the mapping table includes: determining the data type of each element in the mapping table, where the data type includes a string, an integer, a list, and a nested table; determining the format specification of the corresponding column storage file according to the data type, and constructing a Schema field definition based on the format specification; integrating each Schema field definition to generate a data structure.

[0072] S250: compress the original semi-structured data based on the header information and store it in a column storage file.

[0073] Optionally, the original semi-structured data is compressed and stored in a column storage file based on the header information, including: traversing each field in the original semi-structured data and its corresponding field data; taking each field as a target field in turn, and determining the target header information corresponding to the target field in the header information; determining the target column data block corresponding to the target header information from the column storage file; and using a specified compression algorithm to store the target field data corresponding to the target field in the target column data block.

[0074] Optionally, after compressing and storing the original semi-structured data in a columnar storage file based on the header information, the method further includes: determining attribute information of each column data block in the columnar storage file, wherein the attribute information includes a column value range and a row group; and establishing index information for each column data block based on the attribute information.

[0075] Optionally, the method also includes: determining the target data access mode and the target storage device type; obtaining a storage policy list, wherein the storage policy list includes storage policies corresponding to each data access mode and storage device type; matching the target data access mode and the target storage device type through the storage policy list to obtain a corresponding target storage policy, wherein the target storage policy includes data block capacity and compression algorithm type.

[0076] Optionally, the method further includes: performing storage optimization on the columnar storage files according to a specified period, wherein the storage optimization includes merging files, reorganizing data blocks, and updating index information.

[0077] The technical solution of the embodiment of the present invention makes the data structure clearer, easier to understand and process by mapping the key values ​​of each field to construct the element correspondence. The targeted processing method can quickly generate a mapping table and improve the efficiency of the entire parsing process. The management of semi-structured data is made more standardized and orderly, which is convenient for subsequent updates, maintenance and expansion. By making full use of the characteristics of semi-structured data for compression, a better compression effect can be achieved than that of traditional general compression algorithms, saving storage space and reducing the hardware and time costs of enterprises in data storage and processing.

[0078] Embodiment 3

[0079] Figure 3 This is a schematic diagram of the structure of a semi-structured data compression storage device provided in Embodiment 3 of the present invention. Figure 3 As shown, the apparatus includes: a data parsing and mapping module 310, which is used to obtain original semi-structured data and parse the original semi-structured data to generate a mapping table;

[0080] A header information generating module 320 is used to determine the data structure of the column storage file according to the mapping table, and to create header information of the column storage file according to the data structure;

[0081] The data compression storage module 330 is used to compress and store the original semi-structured data into a column storage file based on the header information.

[0082] Optionally, the data parsing and mapping module 310 is specifically used to: traverse each field in the original semi-structured data; perform key-value mapping on each field to construct an element correspondence relationship, and generate a mapping table according to the element correspondence relationship.

[0083] Optionally, the table header information generation module 320 is specifically used to: determine the data type of each element in the mapping table, where the data type includes a string, an integer, a list, and a nested table; determine the format specification of the corresponding column storage file according to the data type, and construct a Schema field definition based on the format specification; integrate the Schema field definitions to generate a data structure.

[0084] Optionally, the data compression storage module 330 is specifically used to: traverse each field in the original semi-structured data and its corresponding field data; take each field as the target field in turn, and determine the target header information corresponding to the target field in the header information; determine the target column data block corresponding to the target header information from the column storage file; and use a specified compression algorithm to store the target field data corresponding to the target field in the target column data block.

[0085] Optionally, the device also includes: an index information establishment module, which is used to determine the attribute information of each column data block in the column storage file after compressing and storing the original semi-structured data in the column storage file based on the header information, wherein the attribute information includes a column value range and a row group; and establish index information for each column data block based on the attribute information.

[0086] Optionally, the device also includes: a target storage policy determination module, which is used to: determine the target data access mode and the target storage device type; obtain a storage policy list, wherein the storage policy list includes storage policies corresponding to each data access mode and storage device type; match the target data access mode and the target storage device type through the storage policy list to obtain the corresponding target storage policy, wherein the target storage policy includes data block capacity and compression algorithm type.

[0087] Optionally, the device further includes: a storage optimization module, used to: perform storage optimization on the column storage files according to a specified period, wherein the storage optimization includes merging files, reorganizing data blocks, and updating index information.

[0088] The technical solution of the embodiment of the present invention, by adopting column storage and a specific compression method, can read the value of a specific field without decompressing the entire data, thereby greatly improving the efficiency of data reading. It avoids the inconvenience of reading caused by overall compression, makes the use of data more flexible and convenient, and improves the readability of data. By making full use of the characteristics of semi-structured data for compression, it can achieve a better compression effect than traditional general compression algorithms, save storage space, and reduce the hardware and time costs of enterprises in data storage and processing.

[0089] A semi-structured data compression storage device provided by an embodiment of the present invention can execute a semi-structured data compression storage method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.

[0090] Embodiment 4

[0091] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0092] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0093] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0094] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a semi-structured data compression storage method.

[0095] In some embodiments, a semi-structured data compression storage method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the semi-structured data compression storage method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform a semi-structured data compression storage method in any other appropriate manner (e.g., by means of firmware).

[0096] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0097] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0098] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0099] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0100] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0101] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0102] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0103] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for compressing and storing semi-structured data, characterized in that: include: Acquire original semi-structured data, and parse the original semi-structured data to generate a mapping table; Determine the data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure; The original semi-structured data is compressed and stored in the column storage file based on the header information.

2. The method according to claim 1, characterized in that The parsing of the original semi-structured data to generate a mapping table includes: Traversing each field in the original semi-structured data; Key-value mapping is performed on each of the fields to construct an element correspondence relationship, and the mapping table is generated according to the element correspondence relationship.

3. The method according to claim 2, characterized in that Determining the data structure of the column storage file according to the mapping table includes: Determine the data type of each element in the mapping table, wherein the data type includes a string, an integer, a list, and a nested table; Determine the format specification of the corresponding column storage file according to the data type, and construct a Schema field definition based on the format specification; The Schema field definitions are integrated to generate the data structure.

4. The method according to claim 1, characterized in that: The compressing and storing the original semi-structured data in the column storage file based on the header information includes: Traversing each field in the original semi-structured data and its corresponding field data; Taking each of the fields as a target field in turn, and determining target header information corresponding to the target field in the header information; Determining a target column data block corresponding to the target header information from the column storage file; The target field data corresponding to the target field is stored in the target column data block by using a specified compression algorithm.

5. The method according to claim 4, characterized in that After compressing and storing the original semi-structured data in the column storage file based on the header information, the method further includes: Determine attribute information of each column data block in the column storage file, wherein the attribute information includes a column value range and a row group; Index information is created for each column data block based on the attribute information.

6. The method according to claim 1, characterized in that The method further comprises: Determine target data access patterns and target storage device types; Obtaining a storage policy list, wherein the storage policy list includes storage policies corresponding to each data access mode and storage device type; The target data access mode and the target storage device type are matched through the storage policy list to obtain a corresponding target storage policy, wherein the target storage policy includes a data block capacity and a compression algorithm type.

7. The method according to claim 5, characterized in that The method further comprises: The column storage file is optimized for storage according to a specified period, wherein the storage optimization includes merging files, reorganizing data blocks, and updating index information.

8. A semi-structured data compression storage device, characterized in that: include: A data parsing and mapping module, used to obtain original semi-structured data, and parse the original semi-structured data to generate a mapping table; A header information generation module, used to determine the data structure of the column storage file according to the mapping table, and create header information of the column storage file according to the data structure; A data compression storage module is used to compress and store the original semi-structured data into the column storage file based on the header information.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Compression storage method for multi-version parameters of semiconductor process formula and related products

    CN120180989A

  • Method and device for fusing, storing and managing massive small files and data records

    CN121833631A