Data management method, device and system
By adopting the same data file structure for data storage in the data management system, the integrated reading and writing of multiple storage interfaces is realized, and compression processing is supported during the data reading and writing process, the problems of data migration and replication in the existing system are solved, and data reading and writing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202311641802.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-03
AI Technical Summary
Existing data management systems cannot realize the integrated reading and writing of multiple protocols, resulting in multiple data replication and migration when calling the same copy of data under different protocols. It consumes too much communication traffic when reading and writing a large amount of data, occupies too much storage space, and has low reading and writing efficiency.
By using the same data file structure to store data in the form of data storage blocks, the integrated reading and writing of multiple storage interfaces can be realized, data migration and copying can be avoided, and data compression processing is supported during data reading and writing, saving communication traffic and storage space.
It realizes that different storage interfaces do not need to migrate or copy when calling the same data, which saves communication traffic and storage space during the data reading and writing process, and improves data reading and writing efficiency.
Smart Images

Figure CN120086280A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly, to a data management method, apparatus, and system. Background Art
[0002] In data services, a piece of data may need to be utilized multiple times in one scenario, and the same piece of data may also be utilized in different scenarios. Since existing data management systems cannot achieve the fusion of read and write of multiple protocols, when the same piece of data is called using different protocols, multiple data copies and migrations are required. At the same time, when existing data management systems read and write a large amount of data, excessive communication traffic will be consumed during the data read and write process, occupying too much data storage space, and the data read and write efficiency is also low. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data management method, apparatus, and system to achieve the fusion of read and write of multiple storage interfaces, so that when different storage interfaces call the same piece of data, there is no need to migrate and copy the data. At the same time, the communication traffic consumed during the data read and write process is saved, the occupation of data storage space is reduced, and the data read and write efficiency is improved.
[0004] In a first aspect, an embodiment of the present invention provides a data management method, the method comprising:
[0005] Receiving a write data request corresponding to a target data file;
[0006] Determining compression requirement information corresponding to the target data file;
[0007] Determining at least one data logical block according to the write data request and the compression requirement information;
[0008] Determining at least one data shard corresponding to each of the data logical blocks;
[0009] Copying the data corresponding to each of the data shards to the corresponding at least one data storage block;
[0010] Storing each of the data storage blocks into a data storage space according to the compression requirement information, and updating metadata, where the metadata includes storage structure information of the target data file.
[0011] In a second aspect, an embodiment of the present invention provides a data management method, the method comprising:
[0012] Receiving a read data request corresponding to a target data file;
[0013] Obtaining the metadata of the target data file, where the metadata includes storage structure information of the target data file;
[0014] Determine a set of required data logical blocks in the target data file according to the meta information;
[0015] Determine a list of target data shards according to the set of data logical blocks;
[0016] Determine compression requirement information corresponding to the target data file;
[0017] Read the data in the required target data storage blocks in the data storage space according to the compression requirement information and the list of target data shards to obtain target read data.
[0018] In a third aspect, an embodiment of the present invention provides a data management system, and the system includes:
[0019] A storage interface service module with multiple storage interfaces, configured to execute the method described in the first aspect or the second aspect;
[0020] A meta information service module, configured to store and manage the meta information of the target data file, where the meta information includes the storage structure information of the target data file.
[0021] In a fourth aspect, an embodiment of the present invention provides an electronic device, and the device includes:
[0022] A memory for storing one or more computer program instructions;
[0023] A processor, where the one or more computer program instructions are executed by the processor to implement the method described in the first aspect or the second aspect.
[0024] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method described in the first aspect or the second aspect.
[0025] Embodiments of the present invention will store data in the form of data storage blocks using the same data file structure in various storage services. Thus, the fusion reading and writing of multiple storage interfaces can be realized, so that when different storage interfaces call the same data, there is no need to migrate and copy the data. At the same time, embodiments of the present invention will support compression processing of data during the data reading and writing process. Thus, the communication traffic consumed during the data reading and writing process can be saved, the occupation of the data storage space can be reduced, and the data reading and writing efficiency can be improved. Description of the Drawings
[0026] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0027] Figure 1 Schematic diagram of the data management system according to an embodiment of the present invention;
[0028] Figure 2 Schematic diagram of the data file information according to an embodiment of the present invention;
[0029] Figure 3 Flowchart of the data management method according to an embodiment of the present invention;
[0030] Figure 4 Schematic diagram of the data storage process according to an embodiment of the present invention;
[0031] Figure 5 Flowchart of the data management method according to an embodiment of the present invention;
[0032] Figure 6 Schematic diagram of the data reading process according to an embodiment of the present invention;
[0033] Figure 7 Schematic diagram of the data management device according to an embodiment of the present invention;
[0034] Figure 8 Schematic diagram of the data management device according to an embodiment of the present invention;
[0035] Figure 9 Schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners
[0036] The following description is based on embodiments of the present application, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, flows, elements and circuits are not described in detail.
[0037] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0038] Unless the context clearly requires otherwise, the words "including", "comprising" and the like in the entire application document should be construed as including rather than exclusive or exhaustive; that is, the meaning of "including but not limited to".
[0039] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0040] In the solutions described in this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.
[0041] An embodiment of the present invention provides a data management system. The data management system can provide a unified data read-write entry for various computing applications in the data lake ecosystem, and integrate multiple storage protocols to support various storage services such as Hadoop storage services (such as storage services based on the HDFS distributed file system and the HBase distributed NoSQL database), S3 (S3 Simple Storage Service), K8S CSI storage service, and Posix (Portable Operating System Interface of UNIX) for read-write operations. Further, the data management system can be applied in various data ecosystems and application scenarios to achieve the integrated read-write of multiple storage interfaces, save the communication traffic consumed during data read-write, reduce the occupation of data storage space, and improve the data read-write efficiency.
[0042] Figure 1 is a schematic diagram of the data management system according to an embodiment of the present invention. As Figure 1 shown, the data management system 12 is respectively connected to the service layer 11 and the storage layer 13.
[0043] Among them, the data management system 12 integrates multiple storage protocols to achieve the integrated read-write of multiple storage interfaces, so that various computing application clients in the service layer 11 can call appropriate storage interfaces to write data into the storage layer 13 or read the required data from the storage layer 13.
[0044] Specifically, various computing application clients in the service layer 11 can initiate a write data request or a read data request to the data management system 12 through any storage interface supported by the data management system 12. When the data management system 12 receives the write data request or the read data request issued by the service layer 11, it can write the data into the corresponding data file in the storage layer 13 or read the required data from the corresponding data file in the storage layer 13 based on the corresponding storage protocol.
[0045] Further, the data management system 12 may include a storage interface service module 121 and a meta-information service module 122 having multiple storage interfaces.
[0046] Optionally, the storage interfaces of the storage interface service module 121 may include interfaces such as Posix file interfaces, HDFS SDK, CSI, S3 interfaces, S2 interfaces, and image processing interfaces. It should be understood that the above storage interfaces are not limited, and other storage interfaces for implementing storage services may also be integrated into this embodiment.
[0047] Furthermore, in this embodiment, regardless of the type of storage protocol (such as S3 or file system, etc.), the data file may include two parts: meta information and data.
[0048] Furthermore, the storage interface service module 121 may be used to split the data file into multiple data logical blocks (Chunks). Optionally, each Chunk of the data file may have a configurable fixed size. For example, the size of each Chunk may be configured to 64M. It should be understood that the size of each Chunk in the data file may be configured according to the data type of the data file, the data application scenario, and / or the parameters of the corresponding data storage space, etc. This embodiment does not limit the size of the Chunk. In other optional implementation manners, the sizes of the respective Chunks of the data file may also be not fixed, and this embodiment does not limit this.
[0049] Furthermore, each data logical block may be composed of multiple data slices (Slices). Optionally, the size of each Slice may be customized, as long as it is less than the size of the corresponding Chunk. It should be understood that the size of the Slice may be customized according to the size of the data to be written and / or the data type, etc. This embodiment does not limit this.
[0050] Furthermore, each data slice may be composed of multiple data storage blocks (Blocks). Optionally, each Block may have a configurable fixed size. For example, the size of each Block may be configured to 4M, etc. It should be understood that the size of each Block may be configured according to the data type of the data file, the data application scenario, and / or the parameters of the corresponding data storage space, etc. In other optional implementation manners, the sizes of the respective Blocks of the data file may also be not fixed, and this embodiment does not limit this.
[0051] Furthermore, after the data in the data file is written into each Block, the data in each Block may be stored by the storage interface service module 121 into the corresponding data storage space (such as GIFT DFS storage system, S3 storage system, OSS storage system, and COS storage system, etc.).
[0052] Further, the meta - information service module 122 is configured to store and manage the meta - information of data files. Optionally, the meta - information can be stored by the meta - information service module 122 into a meta - information service (such as an MDS meta - information service). Further still, the meta - information may include file index information and storage structure information of the data file. Specifically, the storage structure information may be the correspondence relationship between each Chunk, each Slice, and each Block of the data file.
[0053] In this embodiment, the meta - information can be used to assist the storage interface service module 121 in writing the split data into the data file and can also be used to assist the storage interface service module 121 in reading the data from the data file. For example, when writing data S, the storage interface service module 121 can use the meta - information to split data S and write it into the data file. When reading data S, the storage interface service module 121 can use the meta - information to read data S from the data file.
[0054] In an optional implementation manner, the file index information may include the identifier (vid) of the data file, the index identifier (node_id), the data type (type) of the data file, the data length (length) of each field, etc. The storage structure information may include the Chunk table corresponding to the data file. The Chunk table may include the identifier (vid) of the data file, the corresponding Slice table, and the belonging index identifier (node_id), etc. The Slice table may include information such as the identifier (ID) of each Slice, the size (Size) of the Slice, and the offset (Offset).
[0055] Further, after writing the data in the data file into each Block and storing the data in each Block into the data storage space, in order to ensure that new data can continue to be stored or the required data can be read using the meta - information, the meta - information service module 122 can also update the meta - information of the data file.
[0056] Thus, in this embodiment, by storing the data of the data file in the form of data storage blocks into the data storage space and updating the meta - information of the data file after each data storage, it can be ensured that regardless of which protocol is used for data storage, the data file has the same data storage method and meta - information. Furthermore, it is possible to achieve the integration of multiple storage protocols, so that the data file can be read using any supported storage protocol without replication and data migration.
[0057] Further, in this embodiment, before splitting the data to be written into multiple Chunks, the storage interface service module 121 may also perform compression processing on the data to be written according to the compression requirement information. Also, before storing the data in each Block into the data storage space, the storage interface service module 121 may also perform compression processing on the data in each Block according to the compression requirement information.
[0058] Thus, by supporting compression processing of data during the data storage process, this embodiment can ensure that both the data stored in the data storage space and the data read from the data storage space are compressed data, thereby saving the communication traffic consumed during data reading and writing, reducing the occupation of the data storage space, and improving the data reading and writing efficiency.
[0059] In an optional implementation manner, when performing compression processing on the data in each Block and storing it in the data storage space, the storage interface service module 121 may use a corresponding compression algorithm (such as the Snappy compression algorithm) to perform compression processing on the data in each Block, and then store the compressed data in each Block in the data storage space (such as the GIFT DFS storage system, S3 storage system, OSS storage system, and COS storage system, etc.). At this time, the compression processing of the data in each Block is executed within the data management system. Alternatively, in another optional implementation manner, when performing compression processing on the data in each Block and storing it in the data storage space, the storage interface service module 121 may also directly send the data in each Block and the corresponding compression requirement information to the data storage space, so that the data storage space compresses and stores the data in each Block according to the compression requirement information. At this time, the compression processing of the data in each Block can be executed within the data storage space.
[0060] In an optional implementation manner, the compression requirement information may represent the compression requirement for a single data file, or the compression requirement for all data files in the bucket table / volume table where the data file is located. Alternatively, in some embodiments, the compression requirement information may also be determined temporarily according to the data type of the data to be written, and the present application does not limit this.
[0061] In an optional implementation, the data management system 12 may further include a configuration module (not shown in the figure). The configuration module may be configured to store and manage configuration information to configure various types of storage interfaces, thereby realizing fusion reading and writing of various types of storage interfaces. Specifically, this embodiment may configure a unified bucket table / volume table for each storage service, and the configuration information of the bucket table / volume table may be stored in the configuration module. Optionally, the bucket table / volume table may be created based on user operations, and the configuration information may include name information of the bucket table / volume table, creation area information, project name information of the corresponding project, generation time information, storage data type information and / or other related information. It should be understood that when the compression requirement information represents the compression requirement for all data files in the bucket table / volume table where the data file is located, the configuration information may also include the compression requirement information.
[0062] Therefore, by configuring a unified bucket table / volume table for each storage service, this embodiment can connect the data of various storage services, so that data can be interoperable between the storage services. For example, when the POSIX storage service wants to access the bucket / volume under the S3 storage service, since the POSIX storage service and the S3 storage service are configured with a unified bucket / volume, this embodiment can directly mount the bucket / volume under the S3 storage service to the system for access by the POSIX storage service.
[0063] Furthermore, in this embodiment, the meta information service can be deployed in a corresponding server, and the server can be a single computer or a server cluster composed of multiple computers. Correspondingly, the data storage space can also be deployed in a corresponding server, and the server can also be a single computer or a server cluster composed of multiple computers, and this embodiment does not limit this. Furthermore, in this embodiment, the data management device for executing the data management method can be connected to the above two servers by wired or wireless means, so as to realize the data management system in this embodiment.
[0064] Figure 2 Schematic diagram of data file information according to an embodiment of the present invention. Figure 2 As shown, the data file 21 may include two parts: data and meta information. The meta information may include file index information and data structure information.
[0065] Further, in this embodiment, the file index information may include information such as the identifier (vid) of the data file 21, the index identifier (node_id), the data type (type) of the data file, and the data length (length) of each field. The data structure information may include information such as the index information (Inode) of the data file 21, the Chunk table, and the Slice table.
[0066] In this embodiment, the data of the data file 21 can be split into n1 data logical blocks (Chunk1,..., Chunk i1,..., Chunk n1) after being compressed. Among them, 1 ≤ i1 ≤ n1, and n1 ≥ 1. The Chunk table may include information such as the identifier (vid) of the data file 21, the corresponding Slice table (slices), and the belonging index identifier (node_id).
[0067] In this embodiment, each Chunk may further include multiple data slices. For example, Chunk i1 includes n2 data slices (Slice 1,..., Slice i2,..., Slice n2). Among them, 1 ≤ i2 ≤ n2, and n2 ≥ 1. Among them, the Slice table includes the meta-information of each data slice. Further, the meta-information 22 of the data slice Slice includes fields (Field) such as the identifier (ID), the size (Size) of the Slice, the offset (Offset), and the length (Length) of each field. For example, the length of the field ID is uint64, the length of the field Size is uint32, and the length of the field Offset is uint32.
[0068] In this embodiment, each data slice may further include multiple data storage blocks (Block). For example, data slice Slice 1 includes n3 data storage blocks (Block 1,..., Block i3,..., Block n3). Among them, 1 ≤ i3 ≤ n3, and n3 ≥ 1.
[0069] Further, after the data of the data file 21 is written into each Block, the data in each Block can be further compressed and stored in the data storage space 23. The meta-information can be stored in the meta-information service 24.
[0070] Figure 3 It is a flowchart of the data management method according to an embodiment of the present invention. As Figure 3 shown, the data management method according to an embodiment of the present invention may specifically include the following steps:
[0071] It should be understood that the execution subject of the data management method may specifically be the data management system in the above embodiment, by executingFigure 3 The data management method shown, the data management system can achieve the storage of data.
[0072] S100. Receive a write data request corresponding to the target data file.
[0073] Specifically, various computing application clients in the service layer can initiate a write data request to the data management system through any storage interface supported by the data management system, and the data management system can receive the write data requests issued by each client. Among them, the target data file can be the data file for which a data writing operation is to be performed currently.
[0074] Optionally, in step S100, the write data request may include file index information of the target data file. After receiving the write data request, the data management system can determine the corresponding target data file according to the file index information parsed from the write data request.
[0075] Optionally, the write data request may further include data to be written, that is, the data to be written into the target data file. The data management system can parse the write data request to obtain the data to be written.
[0076] S200. Determine compression requirement information corresponding to the target data file.
[0077] Specifically, the data management system can determine compression requirement information corresponding to the target data file. Among them, the compression requirement information can represent the compression requirement.
[0078] Optionally, in this embodiment, what the compression requirement information represents can be the compression requirement for the target data file, or the compression requirement for all data files in the bucket table / volume table where the target data file is located. Further, when the compression requirement information represents the compression requirement for the target data file, the compression requirement information can be set by the user according to actual needs when creating the target data file. In step S200, the data management system can determine the compression requirement information according to the creation information of the target data file. When the compression requirement information represents the compression requirement for all data files in the bucket table / volume table where the target data file is located, the compression requirement information can be configured by the user according to actual needs when creating the bucket table / volume table. In step S200, the data management system can determine the compression requirement information according to the configuration information of the bucket table / volume table where the target data file is located.
[0079] Optionally, in this embodiment, the types of compression algorithms supported by the data management system may include ZSTD (Zstandard), Snappy, GZIP, etc. Among them, ZSTD is an efficient compression algorithm, which can compress data based on the principles of dictionary coding and entropy coding. Snappy is a high-speed compression and decompression algorithm, which has high speed and reasonable compression ratio, so it is often used in many internal systems. GZIP is a data compression algorithm, which can reduce the file size by deleting duplicate data, so as to achieve data compression. It should be understood that other types of compression algorithms, such as brotli and lz4, etc., can also be integrated into the data management system, and this application does not limit this.
[0080] Optionally, in this embodiment, the data management system may also only support compressing the to-be-written data that has not been compressed. Specifically, when determining the compression requirement information, when the data management system detects that the to-be-written data is uncompressed data, it can determine that the compression requirement information is that no first compression and second compression are required. Further, in this embodiment, the data management system can determine whether the current to-be-written data is uncompressed data according to the file suffix of the current to-be-written data. Specifically, when the data management system detects that the file suffix of the current to-be-written data belongs to the preset compressed file suffixes (such as zip, 7z, rar, gz, xz, bz2, tar, tar.gz, tar.xz, and tar.bz2, etc.), it can determine that the current to-be-written data is uncompressed data. After the to-be-written data is determined to be uncompressed data, the compression requirement information of the current to-be-written data will be set to no first compression and second compression required. Furthermore, in the subsequent method process, the current to-be-written data can still be split by the data management system, but will no longer be compressed by the data management system.
[0081] S300. Determine at least one data logical block according to the write data request and the compression requirement information.
[0082] Specifically, after obtaining the compression requirement information, the data management system can determine at least one data logical block according to the write data request and the compression requirement information.
[0083] Optionally, in this embodiment, two data compression processes may be involved in the entire data storage process, namely the first compression and the second compression. The compression requirement information may specifically be used to characterize whether the first compression and the second compression are required, and to characterize the types of compression algorithms required for the first compression and the second compression. In step S300, the data management system may determine whether the current data to be written needs to be first compressed according to the compression requirement information. If the determination result is that the first compression is required, the data management system may perform the first compression on the current data to be written, and then determine each data logical block according to the compressed current data to be written. If the determination result is that the first compression is not required, the data management system may directly determine each data logical block according to the current data to be written.
[0084] Further, when determining each data logical block, if the data logical blocks (Chunks) in the target data file are pre-configured with a fixed size, then the data management system may determine each data logical block according to the data size of the data to be written and the first offset. The first offset may be used to characterize the number of data logical blocks that have been fully used in the target data file (for example, if the first N data logical blocks in the target data file have been fully used, the first offset may be N).
[0085] For example, assume that the size of the data logical blocks in the target data file is configured to be 64M, the data size of the data to be written is 68M, and the first offset is 1. Then the data management system may determine that 2 data logical blocks are required according to the data size of the data to be written, and then determine that the required data logical blocks are the 2nd and 3rd data logical blocks in the data file according to the first offset.
[0086] S400. Determine at least one data shard corresponding to each of the data logical blocks.
[0087] Specifically, after determining each data logical block, the data management system may respectively determine at least one data shard corresponding to each data logical block.
[0088] Optionally, in step S400, if the currently determined data logical block is an unused data logical block, the data management system may directly create a new data shard in the data logical block. If the currently determined data logical block is a partially used data logical block, the data management system may determine whether the data shards in the data logical block are reusable. If the determination result is that they are reusable, the data management system may directly reuse the data shards in the data logical block until the amount of data written reaches the size of the data shard. If the determination result is that they are not reusable, the data management system may create a new data shard in the data logical block.
[0089] Further, the data management system can obtain the target data file context and determine whether a data slice is reusable by judging whether the data to be written is coherent with the data in the data slice according to the target data file context. Specifically, if the judgment result is that the data to be written is coherent with the data in the data slice, the data management system can determine that the data slice is reusable. If the judgment result is that the data to be written is not coherent with the data in the data slice, the data management system can determine that the data slice is not reusable. Thus, this embodiment can reduce the number of data slices in the data logical block by reusing data slices, thereby reducing the complexity of the corresponding meta-information (i.e., reducing the content of the Slice table) and improving data access efficiency.
[0090] For example, assume that the determined data logical block includes Chunk2 and Chunk3, and the data logical block Chunk2 includes data slices Slice2 and Slice3. Further, if the currently written data is not coherent with the data previously written to Slice3, the data slice Slice3 is not reusable. At this time, a new data slice Slice4 can be created in the data logical block Chunk2. Further, assume that the size of the data to be written is greater than the size of the data slice Slice4. Then, a new data slice can be created in the data logical block Chunk3. For example, assume that the size of the data to be written is 68M and the size of the data slice Slice4 is less than 58M. Then, a new data slice Slice5 can be created in the data logical block Chunk3 based on the corresponding offset. Thus, this embodiment can obtain the data slices Slice4 and Slice5 in the data logical blocks Chunk2 and Chunk3.
[0091] It should be understood that whether the currently written data is coherent with the previously written data can be determined by the data management system according to the target data file context.
[0092] S500. Copy the data corresponding to each of the data slices to the corresponding at least one data storage block.
[0093] Specifically, after determining each data slice, the data management system can copy the data corresponding to each data slice to the corresponding at least one data storage block.
[0094] Optionally, in step S500, the data management system may copy the data in each of the data shards to the corresponding data storage block according to the data size of each data shard and the second offset. Wherein, the second offset is used to represent the number of data storage blocks that have been used in the data shard (for example, if the first M data logical blocks in the data shard have been used, the second offset may be M). It should be understood that for a newly created data shard, its second offset may be 0.
[0095] For example, assume that the data sizes of data shards Slice4 and Slice5 in data logical blocks Chunk2 and Chunk3 are 24M and 44M respectively, the second offset of data shard Slice4 is 10, and the second offset of data shard Slice5 is 0. Then for data shard Slice4, the data processing system may first determine that the data of Slice4 requires 6 data storage blocks (each data storage block has a size of 4M) according to the data size of data shard Slice4, and then split and copy the data of data shard Slice4 to the 11th - 16th data storage blocks of Slice4 according to the second offset of data shard Slice4. For data shard Slice5, the data processing system may first determine that the data of Slice5 requires 11 data storage blocks (each data storage block has a size of 4M) according to the data size of data shard Slice5, and then split and copy the data of data shard Slice5 to the 1st - 11th data storage blocks of Slice5 according to the second offset of data shard Slice5.
[0096] S600. Store each of the data storage blocks into the data storage space according to the compression requirement information, and update the meta - information.
[0097] Specifically, after copying the data corresponding to each data shard to the corresponding at least one data storage block, the data management system may store each data storage block into the data storage space according to the compression requirement information, and then update the meta - information. Wherein, the meta - information includes the storage structure information of the target data file.
[0098] Optionally, in step S600, the data management system may determine whether the data in each data storage block needs to be secondarily compressed according to the compression requirement information. If the judgment result is that secondary compression is required, the data management system may perform secondary compression on the data in each data storage block according to the compression requirement information, and then store the compressed data in each data storage block into the data storage space. If the judgment result is that secondary compression is not required, the data management system may directly store the data in each data storage block into the data storage space.
[0099] Optionally, when the judgment result indicates that the second compression is required, the data management system may also directly send the data in each Block and the corresponding compression requirement information to the data storage space, so that the data storage space performs the second compression on the data in each Block according to the compression requirement information and stores it.
[0100] It should be understood that the types of compression algorithms used for the first compression and the second compression in this embodiment may be the same or different, and the present application does not limit this. Further, this embodiment may perform only the first compression, or only the second compression, or may also perform the first compression and the second compression simultaneously, and the present application does not limit this.
[0101] It should be understood that updating the meta information may specifically refer to updating the file index information (such as data type and data file length, etc.) and storage structure information (such as the first offset and the second offset, etc.) of the target data file that has changed after storing the data in each data storage block Block into the data storage space.
[0102] In the embodiments of the present invention, the same data file structure will be adopted in various storage services to store data in the form of data storage blocks. Thus, the fusion reading and writing of multiple storage interfaces can be realized, so that when different storage interfaces call the same piece of data, there is no need to migrate and copy the data. At the same time, the embodiments of the present invention support compressing the data during the data reading and writing process. Thus, the communication traffic consumed during the data reading and writing process can be saved, the occupation of the data storage space can be reduced, and the data reading and writing efficiency can be improved.
[0103] Figure 4 It is a schematic diagram of the data storage process of the embodiments of the present invention. As Figure 4 shown, after receiving a write data request for a target data file, the data management system may parse the write data request to obtain the data to be written 41, and obtain the compression requirement information corresponding to the target data file.
[0104] Further, if the obtained compression requirement information indicates that the first compression is required, the data management system may perform the first compression on the data to be written 41 according to the compression requirement information to obtain the compressed data to be written 42.
[0105] Further, after obtaining the data to be written 42, the data management system may determine the data logical block Chunk2 and the data logical block Chunk3 according to the data size of the data to be written 42 and the first offset.
[0106] Further, after determining data logical blocks Chunk2 and Chunk3, for the partially used data logical block Chunk2, when the data management system detects that the currently to-be-written data is not coherent with the previously written data in data shard S2, it can confirm that data shard S2 cannot be reused. Further, the data management system can create data shard S3 in data logical block Chunk2. For the unused data logical block Chunk3, the data management system can directly create data shard S4 in data logical block Chunk3.
[0107] Further, after determining data shards Slice3 and Slice4, the data management system can write the data corresponding to data shards Slice3 and Slice4 into data storage blocks B3, B4, and B5 according to the second offsets and data sizes of data shards Slice3 and Slice4.
[0108] Further, if the compression requirement information also indicates that a second compression is required, the data management system can perform the second compression on the data in data storage blocks B3, B4, and B5.
[0109] Further, after performing the second compression, the data management system can store the compressed data in data storage blocks B3, B4, and B5 into data storage space 43. At the same time, the data management system can also update the meta information of the target data file. Optionally, the data management system can also send the data in data storage blocks B3, B4, and B5 and the compression requirement information to data storage space 43, so that data storage space 43 performs the second compression on the data in data storage blocks B3, B4, and B5 according to the compression requirement information and stores it.
[0110] Figure 5 This is a flowchart of the data management method according to an embodiment of the present invention. As Figure 5 shown, the data management method according to an embodiment of the present invention may specifically include the following steps:
[0111] It should be understood that the execution subject of the data management method may specifically be the data management system in the above embodiment. By executing Figure 5 the data management method shown, the data management system can realize the reading of the required data.
[0112] S100': Receive a read data request corresponding to a target data file.
[0113] Specifically, various computing application clients in the service layer initiate a read data request to the data management system through any storage interface supported by the data management system, and the data management system can receive the read data request issued by the service layer. Among them, the target data file can be the data file for which a data reading operation is to be performed currently.
[0114] Optionally, in step S100', the read data request may include the file index information of the target data file. After receiving the read data request, the data management system can determine the target data file according to the file index information parsed from the read data request.
[0115] S200': Obtain the meta information of the target data file.
[0116] Specifically, the data management system can obtain the meta information of the target data file. Among them, the meta information includes the storage structure information of the target data file.
[0117] S300': Determine the set of required data logical blocks in the target data file according to the meta information.
[0118] Specifically, after obtaining the meta information, the data management system can determine the set of required data logical blocks in the target data file according to the meta information. Among them, the set of data logical blocks may include at least one target data logical block, and the target data logical block can be the data logical block in the data logical blocks of the target data file corresponding to the data to be read.
[0119] Optionally, the data required to be read by the service layer is not necessarily all the data in the data file. Therefore, the read data request may also include a data reading range (for example, the data to be read can be the data during the period from January 1, 2022 to February 1, 2022). In step S300', the data management system can determine the set of required data logical blocks according to the data reading range. Specifically, the data management system can first determine all the data logical blocks of the target data file according to the meta information, and then determine the set of data logical blocks from all the data logical blocks of the target data file according to the data reading range.
[0120] S400': Determine the target data shard list according to the set of data logical blocks.
[0121] Specifically, after determining the set of data logical blocks, the data management system can determine the target data shard list according to the set of data logical blocks.
[0122] Optionally, in step S400', the data management system may determine a target data shard list according to each target data logic block in the data logic block set. The target data shard list may include at least one target data shard, and the target data shard may be a data shard corresponding to the required read data among the data shards of each target data logic block.
[0123] S500': Determine the compression requirement information corresponding to the target data file.
[0124] Specifically, the data management system may determine the compression requirement information corresponding to the target data file. The acquisition method of the compression requirement information may refer to the above step S200 and will not be elaborated here.
[0125] It should be understood that the compression requirement information may represent the compression requirement for the data to be written. Correspondingly, the compression requirement information may also represent the decompression requirement for the required read data.
[0126] S600': Read the data in the required target data storage block in the data storage space according to the compression requirement information and the target data shard list to obtain the target read data.
[0127] Specifically, after determining the compression requirement information and the target data shard list, the data management system may read the data in the required target data storage block in the data storage space according to the compression requirement information and the target data shard list to obtain the target read data.
[0128] Optionally, in step S600', the data management system may respectively determine the target data storage blocks corresponding to the target data shards in the target data shard list, and then read the data in the target data storage blocks in the data storage space.
[0129] Optionally, since the data stored in the data storage space in the data storage process of this embodiment is all compressed data, in step S600', the data read by the data management system from the data storage space is also compressed data. The data management system needs to decompress the read data according to the compression requirement information to obtain the target read data. Specifically, in step S600', the data management system can first read the data in each target data storage block from the data storage space according to the target data shard list, and then determine whether the first decompression is required according to the compression requirement information (the first decompression corresponds to the second compression in the data storage process). If the judgment result is that the first decompression is required, the data management system can perform the first decompression on the data in each target data storage block according to the compression requirement information, and then splice the data in each target data storage block. If the judgment result is that the first decompression is not required, the data management system can directly splice the data in each target data storage block.
[0130] Optionally, when the judgment result is that the first decompression is required, the data management system can also send the corresponding compression requirement information to the data storage space, so that the data storage space performs the first decompression on the data in each Block according to the compression requirement information. After the data storage space completes the decompression process of the data in each Block, the data management system can directly read the decompressed data in each Block from the data storage space.
[0131] Further, after splicing the data in each target data storage block, the data management system can also determine whether the second decompression is required according to the compression requirement information (the second decompression corresponds to the first compression in the data storage process). If the judgment result is that the second decompression is required, the data management system can perform the second decompression on the spliced data according to the compression requirement information to obtain the target read data. If the judgment result is that the second decompression is not required, the data management system can directly determine the spliced data as the target read data.
[0132] Figure 6 is a schematic diagram of the data reading process of the embodiment of the present invention. As Figure 6 shown, after receiving the read data request for the target data file, the data management system can obtain the meta information of the target data file. After obtaining the meta information, the data management system can determine the target data logical blocks Chunk2 and Chunk3 corresponding to the data to be read in the target data file according to the meta information and the read data request, so as to determine the data logical block set.
[0133] Further, after determining the set of data logical blocks, the data management system may determine a target data shard S3 corresponding to the data to be read in the data shards of Chunk2, and determine a target data shard S4 corresponding to the data to be read in the data shards of Chunk3, so as to obtain a target data shard list.
[0134] Further, after obtaining the target data shard list, the data management system may read the data in each target storage data block (i.e., storage data blocks B3 - B5) corresponding to each target data shard in the data storage space according to the target data shard list.
[0135] Further, after reading the data in each target storage data block, the data management system may determine whether to perform first - stage decompression according to the compression requirement information. If first - stage decompression is required, the data management system may perform first - stage decompression on the data in each target storage data block respectively, and then splice the data in the decompressed target storage data blocks together to obtain spliced data 61.
[0136] Optionally, before reading the data in each target storage data block from the data storage space, the data management system may also first determine whether to perform first - stage decompression according to the compression requirement information. If the determination result is that first - stage decompression is required, the data management system may send the corresponding compression requirement information to the data storage space so that the data storage space performs first - stage decompression on the data in each Block according to the compression requirement information. After the data storage space completes the decompression process of the data in each Block, the data management system may directly read the decompressed data in each Block from the data storage space.
[0137] Further, after obtaining the spliced data 61, the data management system may continue to determine whether to perform second - stage decompression on the spliced data 61 according to the compression requirement information. If second - stage decompression is required, the data management system may perform second - stage decompression on the spliced data 61 to obtain the target read data 62. It should be understood that if first - stage decompression is not required, the data management system may directly determine the spliced data 61 as the target read data 62.
[0138] In the embodiments of the present invention, the same data file structure is adopted in various storage services to store data in the form of data storage blocks. Thus, the integrated reading and writing of multiple storage interfaces can be realized, so that when different storage interfaces call the same data, there is no need to migrate and copy the data. At the same time, the embodiments of the present invention support data compression processing during the data reading and writing process. Thus, the communication traffic consumed during the data reading and writing process can be saved, the occupation of the data storage space can be reduced, and the data reading and writing efficiency can be improved.
[0139] Figure 7 Schematic diagram of the data management device according to an embodiment of the present invention. As Figure 7 shown, the data management device according to an embodiment of the present invention includes a first receiving unit 71, a first compression requirement determining unit 72, a first data logical block determining unit 73, a first data sharding determining unit 74, a data storage block determining unit 75, and a storage unit 76.
[0140] Specifically, the first receiving unit 71 is configured to receive a write data request corresponding to a target data file;
[0141] The first compression requirement determining unit 72 is configured to determine compression requirement information corresponding to the target data file;
[0142] The first data logical block determining unit 73 is configured to determine at least one data logical block according to the write data request and the compression requirement information;
[0143] The first data sharding determining unit 74 is configured to determine at least one data shard corresponding to each of the data logical blocks;
[0144] The data storage block determining unit 75 is configured to copy the data corresponding to each of the data shards to at least one corresponding data storage block;
[0145] The storage unit 76 is configured to store each of the data storage blocks into a data storage space according to the compression requirement information and update meta-information, where the meta-information includes storage structure information of the target data file.
[0146] In the embodiment of the present invention, the same data file structure will be adopted in various storage services to store data in the form of data storage blocks. Thus, fusion reading and writing of multiple storage interfaces can be realized, so that when different storage interfaces call the same data, there is no need to migrate and copy the data. At the same time, the embodiment of the present invention supports compression processing of data during data reading and writing. Thus, communication traffic consumed during data reading and writing can be saved, the occupation of data storage space can be reduced, and data reading and writing efficiency can be improved.
[0147] Figure 8 Schematic diagram of the data management device according to an embodiment of the present invention. As Figure 8 shown, the data management device according to an embodiment of the present invention includes a second receiving unit 81, a meta-information obtaining unit 82, a second data logical block determining unit 83, a second data sharding determining unit 84, a second compression requirement determining unit 85, and a reading unit 86.
[0148] Specifically, the second receiving unit 81 is configured to receive a read data request corresponding to a target data file;
[0149] The meta - information acquisition unit 82 is configured to acquire the meta - information of the target data file, where the meta - information includes the storage structure information of the target data file;
[0150] The second data logic block determination unit 83 is configured to determine a set of required data logic blocks in the target data file according to the meta - information;
[0151] The second data sharding determination unit 84 is configured to determine a target data shard list according to the set of data logic blocks;
[0152] The second compression requirement determination unit 85 is configured to determine compression requirement information corresponding to the target data file;
[0153] The reading unit 86 is configured to read the data in the required target data storage blocks in the data storage space according to the compression requirement information and the target data shard list to obtain target read data.
[0154] In the embodiments of the present invention, the same data file structure is adopted in various storage services to store data in the form of data storage blocks. Thus, the fusion reading and writing of multiple storage interfaces can be realized, so that when different storage interfaces call the same data, there is no need to migrate and copy the data. At the same time, the embodiments of the present invention support data compression processing during the data reading and writing process. Thus, the communication traffic consumed during the data reading and writing process can be saved, the occupation of the data storage space can be reduced, and the data reading and writing efficiency can be improved.
[0155] Figure 9 It is a schematic diagram of an electronic device according to an embodiment of the present invention. As Figure 9 shown, Figure 9 The electronic device shown is a general - purpose data processing device, which includes a general - purpose computer hardware structure and at least includes a processor 91 and a memory 92. The processor 91 and the memory 92 are connected through a bus 93. The memory 92 is suitable for storing instructions or programs executable by the processor 91. The processor 91 can be an independent microprocessor or a set of one or more microprocessors. Thus, by executing the instructions stored in the memory 92, the processor 91 executes the method flow of the embodiments of the present invention as described above to implement the processing of data and the control of other devices. The bus 93 connects the above - mentioned multiple components together and at the same time connects the above - mentioned components to a display controller 94, a display device, and an input / output (I / O) device 95. The input / output (I / O) device 95 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well - known in the art. Typically, the input / output (I / O) device 95 is connected to the system through an input / output (I / O) controller 96.
[0156] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus (device), or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0157] The present application is described with reference to the flowcharts of methods, apparatuses (devices), and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.
[0158] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the process Figure 1 specified functions in one or more of the processes.
[0159] These computer program instructions can also be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the Figure 1 specified functions in one or more of the processes.
[0160] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used for a computer to execute some or all of the above method embodiments.
[0161] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks, etc., which can store program codes.
[0162] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data management method, characterized in that, the method includes: receiving a write data request corresponding to a target data file; determining compression requirement information corresponding to the target data file; determining at least one data logical block according to the write data request and the compression requirement information; determining at least one data shard corresponding to each of the data logical blocks; copying the data corresponding to each of the data shards to at least one corresponding data storage block; storing each of the data storage blocks in a data storage space according to the compression requirement information, and updating metadata, where the metadata includes storage structure information of the target data file.
2. The method according to claim 1, characterized in that, the compression requirement information is used to represent whether first compression is required and to represent the type of compression algorithm required for the first compression; determining at least one data logical block according to the write data request and the compression requirement information includes: in response to the compression requirement information indicating that first compression is required, performing first compression on the data to be written according to the compression requirement information; determining each of the data logical blocks according to the data size of the data to be written and a first offset; wherein the first offset is used to represent the number of data logical blocks that have been fully used in the target data file.
3. The method according to claim 1, characterized in that, the compression requirement information is used to represent whether second compression is required and to represent the type of compression algorithm required for the second compression; storing each of the data storage blocks in a data storage space according to the compression requirement information includes: in response to the compression requirement information indicating that second compression is required, performing second compression on the data in each of the data storage blocks according to the compression requirement information; storing the data in each of the data storage blocks in the data storage space.
4. The method according to any one of claims 1-3, characterized in that, determining compression requirement information corresponding to the target data file includes: determining the compression requirement information according to the creation information of the target data file; or determining the compression requirement information according to the configuration information of the bucket table / volume table where the target data file is located.
5. The method according to claim 1, characterized in that, determining at least one data shard corresponding to each of the data logical blocks includes: in response to the data shards in the data logical block being reusable, determining each of the data shards according to the reusable data shards and the data size of the data to be written; in response to the data shards in the data logical block not being reusable, creating new data shards.
6. The method according to claim 5, characterized in that, copying the data corresponding to each of the data shards to at least one corresponding data storage block includes: copying the data in each of the data shards to the corresponding data storage block according to the data size of each of the data shards and a second offset; wherein the second offset is used to represent the number of data storage blocks that have been used in the data shard.
7. The method according to claim 1, characterized in that, Determining the compression requirement information corresponding to the target data file includes: In response to the data to be written being compressed data, determining that the compression requirement information is that no first compression and second compression are required.
8. A data management method, Characterized in that, The method includes: Receiving a read data request corresponding to a target data file; Obtaining the meta information of the target data file, where the meta information includes the storage structure information of the target data file; Determining a set of required data logical blocks in the target data file according to the meta information; Determining a target data shard list according to the set of data logical blocks; Determining the compression requirement information corresponding to the target data file; Reading the data in the required target data storage blocks in the data storage space according to the compression requirement information and the target data shard list to obtain target read data.
9. The method according to claim 8, Characterized in that, The compression requirement information is used to characterize whether first decompression and second decompression are required, and to characterize the types of decompression algorithms required for first decompression and second decompression; Reading the data in the required target data storage blocks in the data storage space according to the compression requirement information and the target data shard list to obtain target read data includes: Reading the data in each of the target data storage blocks in the data storage space according to the target data shard list; In response to the compression requirement information indicating that first decompression is required, performing first decompression on the data in each of the target data storage blocks according to the compression requirement information; Concatenating the data in each of the target data storage blocks; In response to the compression requirement information indicating that second decompression is required, performing second decompression on the concatenated data according to the compression requirement information to obtain the target read data.
10. The method according to claim 8, Characterized in that, The read data request further includes a data reading range; Determining a set of required data logical blocks in the target data file according to the meta information includes: Determining the data logical blocks of the target data file according to the meta information; Determining the set of data logical blocks from the data logical blocks of the target data file according to the data reading range.
11. The method according to claim 8, Characterized in that, The set of data logical blocks includes at least one target data logical block, and the target data logical block is the data logical block in the data logical blocks of the target data file corresponding to the data to be read; Determining a target data shard list according to the set of data logical blocks includes: Determining the target data shard list according to each of the target data logical blocks; Wherein, the target data shard list includes at least one target data shard, and the target data shard is the data shard in the data shards of each of the target data logical blocks corresponding to the data to be read.
12. A data management system, Characterized in that, The system includes: A storage interface service module with multiple storage interfaces, configured to execute the method according to any one of claims 1-11. A meta - information service module, configured to store and manage the meta - information of a target data file, where the meta - information includes the storage structure information of the target data file.
13. The data management system according to claim 12, wherein, the system further includes: a configuration module, configured to generate configuration information for configuring each of the storage interfaces, and the configuration information at least includes a unified bucket / volume list for each of the storage interfaces.
14. An electronic device, wherein, the device includes: a memory, configured to store one or more computer program instructions; a processor, and the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 - 11.
15. A computer - readable storage medium, on which computer program instructions are stored, wherein, the computer program instructions, when executed by a processor, implement the method according to any one of claims 1 - 11.