Data storage method, device, electronic device and storage medium

By obtaining and analyzing the metadata of the target file, generating the interval of the data block, and deleting the part covering the data block when necessary, the resource waste problem during the storage of small files in distributed storage systems is solved, and the effectiveness of data storage is improved.

CN115686378BActive Publication Date: 2025-05-06CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211454534.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-05-06
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

In a distributed storage system, when a server stores data in a data block of the same size, if the amount of data stored is much smaller than the size of the data block, it may lead to waste of disk space resources and affect the effectiveness of data storage.

Method used

By obtaining the metadata of the target file, including the starting offset and size of each data block, the interval for each data block is generated. When the starting offset of the data block to be inserted is within the interval of a certain data block, delete the part covering the data block and flexibly update the size of the data block.

Benefits of technology

It avoids the waste of disk space resources, improves the effectiveness of data storage, and optimizes the utilization of storage resources by flexibly managing the size of data blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115686378B_ABST
    Figure CN115686378B_ABST
Patent Text Reader

Abstract

The present invention provides a data storage method, device, electronic device and storage medium, which relates to the field of computer technology, and solves the technical problem in the related technology that when a server stores data in data blocks of the same size, when the amount of stored data is much smaller than the size of the data block, it may cause a waste of disk space resources and affect the effectiveness of data storage. The method includes: obtaining metadata of a target file; based on the starting offset of each data block in the target file and the size of each data block, generating an interval of each data block, wherein the minimum value in the interval of a data block is the starting offset of the data block in the target file, and the maximum value in the interval of the data block is the ending offset of the data block in the target file; when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, deleting the first overwriting data block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data storage method, device, electronic equipment and storage medium. Background Art

[0002] Currently, in a distributed storage system, a server can store data in each file in data blocks of the same size. For example, in a distributed file system (Hadoop distributed file system, HDFS), the default size of a file's data block is 128 megabits (Mbit).

[0003] However, the above method may not be applicable to the scenario of small files. When the amount of stored data is much smaller than the size of the data block, it may cause a waste of disk space resources and affect the effectiveness of data storage. Summary of the invention

[0004] The present invention provides a data storage method, device, electronic device and storage medium, which solves the technical problem in the related art that a server stores data in data blocks of the same size. When the amount of stored data is much smaller than the size of the data block, it may cause a waste of disk space resources and affect the effectiveness of data storage.

[0005] In a first aspect, the present invention provides a data storage method, comprising: obtaining metadata of a target file, the metadata comprising a starting offset of each data block in at least one data block in the target file and a size of each data block, the at least one data block being a data block included in the target file; generating an interval of each data block based on the starting offset of each data block in the target file and the size of each data block, wherein a minimum value in an interval of a data block is a starting offset of the data block in the target file, a maximum value in an interval of the data block is an ending offset of the data block in the target file, and a starting offset of the data block in the target file is a maximum value in an interval of the data block. The displacement is the distance between the starting position of the data block and the starting position of the target file, and the ending offset of the data block in the target file is the distance between the ending position of the data block and the starting position of the target file; when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, delete the first overlay data block, the first data block is one of the at least one data block, the starting offset of the first overlay data block in the target file is the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file is the ending offset of the first data block in the target file.

[0006] Optionally, the metadata further includes an identifier of a data file space to which each data block belongs and an offset of each data block in the data file space to which each data block belongs; the data storage method further includes: determining whether a starting offset of the first data block in the target file is the same as a first value, and whether an offset of the first data block in the first data file space is the same as a second value, the first value being the sum of a starting offset of the second data block in the target file and a size of the second data block, the second value being the sum of an offset of the second data block in the first data file space and a size of the second data block, the first data file space being the first data file space. a data block and a data file space to which the second data block belongs, the second data block being a previous data block of the first data block in the at least one data block; when the starting offset of the first data block in the target file is the same as the first value, and the offset of the first data block in the first data file space is the same as the second value, the first data block and the second data block are merged to generate a target data block, the ending offset of the target data block in the target file is the same as the ending offset of the first data block in the target file, and the starting offset of the target data block in the target file is the same as the starting offset of the second data block in the target file.

[0007] Optionally, the above-mentioned data storage method also includes: generating an initial interval tree based on the metadata of the target file, the initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block; inserting the data block to be inserted into the initial interval tree to obtain a target interval tree, the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node, the first node is used to represent the first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

[0008] Optionally, the data storage method further includes: obtaining a file identifier, a preset start offset and a preset size of the target file; based on the file identifier, the preset start offset and the preset size of the target file, obtaining a target data block set, the target data block set including M data blocks, the M data blocks belonging to the target file, the start offset of the first data block of the M data blocks in the target file being the same as the preset start offset, the end offset of the last data block of the M data blocks in the target file being the same as a third value, the third value being the sum of the preset start offset and the preset size, M≥1.

[0009] In a second aspect, the present invention provides a data storage device, including: an acquisition module, a processing module and a deletion module; the acquisition module is used to acquire metadata of a target file, the metadata including a starting offset of each data block in at least one data block in the target file and a size of each data block, the at least one data block being a data block included in the target file; the processing module is used to generate an interval of each data block based on the starting offset of each data block in the target file and the size of each data block, wherein the minimum value in the interval of a data block is the starting offset of the data block in the target file, the maximum value in the interval of the data block is the ending offset of the data block in the target file, and the data The starting offset of the block in the target file is the distance between the starting position of the data block and the starting position of the target file, and the ending offset of the data block in the target file is the distance between the ending position of the data block and the starting position of the target file; the deletion module is used to delete the first overlay data block when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, the first data block is one of the at least one data block, the starting offset of the first overlay data block in the target file is the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file is the ending offset of the first data block in the target file.

[0010] Optionally, the metadata further includes an identifier of a data file space to which each data block belongs and an offset of each data block in the data file space to which each data block belongs; the data storage device further includes a determination module; the determination module is used to determine whether a starting offset of the first data block in the target file is the same as a first value, and whether an offset of the first data block in the first data file space is the same as a second value, the first value being the sum of a starting offset of the second data block in the target file and a size of the second data block, the second value being the sum of an offset of the second data block in the first data file space and a size of the second data block, the first data file space being the first value. The processing module is further configured to merge the first data block and the second data block to generate a target data block, wherein the end offset of the target data block in the target file is the same as the end offset of the first data block in the target file, and the start offset of the target data block in the target file is the same as the start offset of the second data block in the target file.

[0011] Optionally, the processing module is also used to generate an initial interval tree based on the metadata of the target file, the initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block; the processing module is also used to insert the data block to be inserted into the initial interval tree to obtain a target interval tree, the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node, the first node is used to represent the first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

[0012] Optionally, the acquisition module is further used to obtain the file identifier, preset start offset and preset size of the target file; the acquisition module is also used to obtain a target data block set based on the file identifier, the preset start offset and the preset size of the target file, the target data block set includes M data blocks, the M data blocks belong to the target file, the start offset of the first data block of the M data blocks in the target file is the same as the preset start offset, the end offset of the last data block of the M data blocks in the target file is the same as a third value, the third value is the sum of the preset start offset and the preset size, and M≥1.

[0013] In a third aspect, the present invention provides an electronic device comprising: a processor and a memory configured to store processor executable instructions; wherein the processor is configured to execute the instructions to implement any one of the optional data storage methods in the first aspect above.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions in the computer-readable storage medium are executed by an electronic device, the electronic device is enabled to execute any one of the optional data storage methods in the first aspect.

[0015] The data storage method, device, electronic device and storage medium provided by the present invention can obtain the metadata of the target file by the terminal, and generate the interval of each data block based on the starting offset of each data block in the target file and the size of each data block in at least one data block included in the metadata. In the case where the starting offset of the data block to be inserted in the target file is within the interval of the first data block (i.e., a data block included in the at least one data block), it means that the data block to be inserted has covered a part of the data blocks in the first data block (i.e., the first covered data block). At this time, the terminal can delete the first covered data block, can flexibly update the size of the data block, avoid the waste of disk space resources, and can improve the effectiveness of data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0017] Figure 1 A schematic diagram of a network architecture of a data storage system provided by an embodiment of the present invention;

[0018] Figure 2 A schematic diagram of a data storage method provided by an embodiment of the present invention;

[0019] Figure 3 A schematic diagram of a situation where there is overlap between data blocks provided by an embodiment of the present invention;

[0020] Figure 4 A schematic diagram of a flow chart of another data storage method provided by an embodiment of the present invention;

[0021] Figure 5 A schematic diagram of a situation where data blocks are merged provided in an embodiment of the present invention;

[0022] Figure 6 A schematic diagram of a flow chart of another data storage method provided by an embodiment of the present invention;

[0023] Figure 7 A schematic diagram of an interval tree provided by an embodiment of the present invention;

[0024] Figure 8 A schematic diagram of a flow chart of another data storage method provided by an embodiment of the present invention;

[0025] Fig. 9 A schematic diagram of the structure of a data storage device provided by an embodiment of the present invention;

[0026] Fig.10 A schematic diagram of the structure of another data storage device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The data storage method, device, electronic device and storage medium provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0028] The terms "first" and "second" in the specification and drawings of this application are used to distinguish different objects rather than to describe a specific order of the objects. For example, the first data block and the second data block are used to distinguish different data blocks rather than to describe a specific order of the data blocks.

[0029] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0030] It should be noted that, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0031] The term "and / or" used in the present application includes using either or both of the two methods.

[0032] In the description of the present application, unless otherwise specified, “plurality” means two or more.

[0033] Based on the description in the background technology, since in the related technology, the server stores data in data blocks of the same size, when the amount of data stored is much smaller than the size of the data block, it may cause a waste of disk space resources, affecting the effectiveness of data storage. Based on this, the embodiment of the present invention provides a data storage method, device, electronic device and storage medium. When the starting offset of the data block to be inserted in the target file is within the interval of the first data block (i.e., a data block included in the at least one data block), it means that the data block to be inserted has covered a part of the data blocks in the first data block (i.e., the first covered data block). At this time, the terminal can delete the first covered data block, can flexibly update the size of the data block, avoid the waste of disk space resources, and can improve the effectiveness of data storage.

[0034] A data storage method, device, electronic device and storage medium provided by an embodiment of the present invention can be applied to a data storage system, such as Figure 1 As shown, the data storage system includes a terminal 101 and a server 102. Generally, in practical applications, the connection between the above-mentioned devices or service functions can be a wireless connection. In order to conveniently and intuitively represent the connection relationship between the various devices, Figure 1 Solid lines are used to indicate.

[0035] The terminal 101 is used to generate an interval of each data block based on the starting offset of each data block in the target file and the size of each data block in at least one data block. The minimum value in the interval of a data block is the starting offset of the data block in the target file, and the maximum value in the interval of the data block is the ending offset of the data block in the target file.

[0036] The server 102 is used to receive a data block acquisition request sent by the terminal 101, the data block acquisition request includes a file identifier of a target file, a preset start offset and a preset size, and the data block acquisition request is used to request to acquire M data blocks, M≥1.

[0037] It should be noted that the terminal 101 may be a mobile phone, tablet computer, desktop, laptop, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR)\virtual reality (VR) device, and the embodiment of the present invention does not impose any special restrictions on the specific form of the terminal 101. It can interact with the user through one or more methods such as keyboard, touch pad, touch screen, remote control, voice interaction or handwriting device.

[0038] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (content delivery network, CDN), as well as big data and artificial intelligence platforms.

[0039] The electronic device for executing the data storage method provided by the embodiment of the present invention is as follows: Figure 1 The terminal 101 shown in the figure is taken as an example to illustrate the data storage method provided by the embodiment of the present invention.

[0040] like Figure 2 As shown, the data storage method provided by the embodiment of the present invention may include S101-S103.

[0041] S101: A terminal obtains metadata of a target file.

[0042] The metadata includes a start offset of each data block in the target file and a size of each data block in at least one data block, and the at least one data block is a data block included in the target file.

[0043] It should be understood that each data block in the at least one data block is used to store data in the target file. For a data block in the at least one data block, the start offset of the data block in the target file is the distance between the start position of the data block and the start position of the target file.

[0044] In the embodiment of the present invention, the starting offset of a data block in the target file may also be understood as the offset (file offset) of the data block in the target file.

[0045] In one implementation of the embodiment of the present invention, the above Figure 1 The server 102 shown in the figure can be understood as a metadata management server, which is used to manage (or store) metadata of each file in multiple files. The terminal can send a metadata acquisition request to the server, and the metadata acquisition request includes the identifier (or name) of the above-mentioned target file, and then the terminal can obtain the metadata of the target file from the server.

[0046] S102: The terminal generates an interval of each data block based on the start offset of each data block in the target file and the size of each data block.

[0047] The minimum value in the interval of a data block is the starting offset of the data block in the target file, and the maximum value in the interval of the data block is the ending offset of the data block in the target file.

[0048] It should be understood that the end offset (end) of a data block in the target file is the distance between the end position of the data block and the start position of the target file.

[0049] In the embodiment of the present invention, for a data block, the terminal may determine the sum of the start offset of the data block in the target file and the size of the data block as the end offset of the data block in the target file, so as to obtain the range of the data block. That is, the terminal may generate the range of each data block.

[0050] S103: When the starting offset of the data block to be inserted in the target file is within the interval of the first data block, the terminal deletes the first overwriting data block.

[0051] Among them, the first data block is one of the at least one data block mentioned above, the starting offset of the first overlay data block in the target file is the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file is the ending offset of the first data block in the target file.

[0052] It should be understood that, when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, it means that the starting offset of the data block to be inserted in the target file is greater than the starting offset of the first data block in the target file, and the starting offset of the data block to be inserted in the target file is less than the ending offset of the first data block in the target file, that is, the data block to be inserted has covered a part of the data block in the first data block (that is, the first covered data block). At this time, the terminal can delete the first covered data block, which can effectively release disk space.

[0053] It can be understood that the first covered data block belongs to the first data block, the first data block also includes a first uncovered data block (i.e., a portion of the data block in the first data block that is not covered by the data block to be inserted), the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file. The terminal can update the first data block to the first uncovered data block, specifically by updating the interval of the first data block to the interval of the first uncovered data block.

[0054] Optionally, the terminal may have a garbage collection (GC) function, and the terminal may delete the first overlay data block based on the GC function.

[0055] In one case, when the end offset of the data block to be inserted in the target file is within the interval of the third data block, the terminal can also delete the third overlay data block, which is a data block other than the first data block in the at least one data block mentioned above, and the start offset of the third overlay data block in the target file is the start offset of the third data block in the target file, and the end offset of the third overlay data block in the target file is the end offset of the data block to be inserted in the target file.

[0056] In another case, when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, and the ending offset of the data block to be inserted in the target file is within the interval of the third data block, the terminal can delete the first overlay data block and the third overlay data block.

[0057] For example, Figure 3 As shown, it is assumed that data block 1 is the first data block, data block 2 is the third data block, and data block 3 is the data block to be inserted. Data block 1 includes data block a and data block b, and data block 2 includes data block c and data block d.

[0058] Assuming that the starting offset of the data block b in the target file is the same as the starting offset of the data block 3 in the target file, and the ending offset of the data block c in the target file is the same as the ending offset of the data block 3 in the target file, the terminal can determine that the data block b is the first overlay data block and the data block c is the third overlay data block. That is, the terminal can delete the data block b and the data block c.

[0059] The technical solution provided by the above embodiment can at least bring the following beneficial effects: From S101-S103, it can be known that the terminal can obtain the metadata of the target file, and based on the starting offset of each data block in the target file and the size of each data block in at least one data block included in the metadata, generate the interval of each data block. In the case where the starting offset of the data block to be inserted in the target file is within the interval of the first data block (i.e., a data block included in the at least one data block), it means that the data block to be inserted has covered a part of the data blocks in the first data block (i.e., the first covered data block). At this time, the terminal can delete the first covered data block, can flexibly update the size of the data block, avoid the waste of disk space resources, and can improve the effectiveness of data storage.

[0060] In an implementation of the embodiment of the present invention, the metadata of the target file further includes an identifier of a data file space to which each data block in at least one data block belongs and an offset of each data block in the data file space to which each data block belongs. Figure 2 ,like Figure 4 As shown, the data storage method provided by the embodiment of the present invention also includes S104-S105.

[0061] S104: The terminal determines whether the start offset of the first data block in the target file is the same as the first value, and whether the offset of the first data block in the first data file space is the same as the second value.

[0062] Among them, the first value is the sum of the starting offset of the second data block in the target file and the size of the second data block, the second value is the sum of the offset of the second data block in the first data file space and the size of the second data block, the first data file space is the data file space to which the first data block and the second data block belong, and the second data block is the previous data block of the first data block in the above-mentioned at least one data block.

[0063] It should be understood that the terminal can determine the data file space to which each data block belongs based on the identifier (segment ID) of the data file space (segment) to which each data block belongs. When a data file space (e.g., a first data file space) is a data file space described by a data block (e.g., a first data block), it means that the first data file space includes the first data block.

[0064] In the embodiment of the present invention, the offset of a data block in the data file space to which the data block belongs can also be understood as the starting offset of the data block in the data file space to which the data block belongs. The offset of the data block in the data file space to which the data block belongs is the distance between the starting position of the data block and the starting position of the data block in the data file space to which the data block belongs.

[0065] In an optional implementation, the metadata management server may store and maintain two data tables, including a file data table and a data file space data table. The file data table includes the starting offset of each data block in the target file, the size of each data block, the identifier of the data file space to which each data block belongs, and the offset of each data block in the data file space to which each data block belongs. The data file space data table includes the identifier of each data file space in at least one data file space, the offset of each data block in the data file space to which each data block belongs, and the size of each data block.

[0066] S105. When the starting offset of the first data block in the target file is the same as the first value, and the offset of the first data block in the first data file space is the same as the second value, the terminal merges the first data block and the second data block to generate a target data block.

[0067] The end offset of the target data block in the target file is the same as the end offset of the first data block in the target file, and the start offset of the target data block in the target file is the same as the start offset of the second data block in the target file.

[0068] In combination with the description of the above embodiment, it should be understood that the first data block and the second data block are both data blocks included in the same file (i.e., the target file), the second data block is the previous data block of the first data block, and the first data block and the second data block are both data blocks included in the same data file space (i.e., the first data file space). When the starting offset of the first data block in the target file is the same as the first value, and the offset of the first data block in the first data file space is the same as the second value, it means that the first data block and the second data block are continuous data blocks, and at this time, the terminal can merge the first data block and the second data block to obtain the target data block. That is, the terminal can merge continuous data blocks to reduce the management resources of data blocks.

[0069] For example, Figure 5 As shown, it is assumed that data block 1 is the first data block mentioned above, and data block 4 is the second data block mentioned above.

[0070] Assume that the first value is the sum of the start offset of data block 4 in the target file and the size of data block 4, and the first value is the same as the start offset of data block 1 in the target file, and the second value is the sum of the offset of data block 4 in the first data file space and the size of data block 4, and the second value is the same as the offset of data block 1 in the first data file space. At this time, the terminal can merge data block 1 and data block 4 to generate Figure 5 The data block 5 shown in the figure is the target data block. The starting offset of the data block 5 in the target file is the same as the starting offset of the data block 4 in the target file, and the ending offset of the data block 5 in the target file is the same as the ending offset of the data block 1 in the target file.

[0071] Combination Figure 2 ,like Figure 6 As shown, the data storage method provided by the embodiment of the present invention may further include S106-S107.

[0072] S106: The terminal generates an initial interval tree based on the metadata of the target file.

[0073] The initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block.

[0074] It should be understood that one node corresponds to one data block, that is, one of the at least one node is used to represent one of the at least one data block. For one of the at least one node, the node may include an interval of the data block represented (or corresponding) by the node.

[0075] Optionally, the node may further include a maximum value (max end), which is the largest value among the maximum value of the interval of the data block represented by the node and the maximum value of the interval of the data blocks represented by all child nodes of the node.

[0076] S107: The terminal inserts the data block to be inserted into the initial interval tree to obtain a target interval tree.

[0077] Among them, the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node mentioned above, the first node is used to represent the above-mentioned first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

[0078] It can be understood that the target interval tree may further include a preset node, where the preset node is used to represent the data block to be inserted, and the preset node includes the interval of the data block to be inserted.

[0079] In combination with the description of the above embodiment, it should be understood that the first uncovered data block is a partial data block in the first data block that is not covered by the data block to be inserted. The terminal can update the first data block, specifically updating the interval of the first data block to the interval of the first uncovered data block. That is, the first node included in the initial interval tree includes the interval of the first data block, and the first node included in the target interval tree includes the interval of the first uncovered data block.

[0080] In one implementation of the embodiment of the present invention, for each node in at least one node (or each data block in at least one data block), the terminal can store the relevant information of each data block in the form of key-value. Specifically, for a node (or a data block), the key of the data block can be the minimum value of the interval of the data block; the value of the data block is the relevant information of the data block, and the relevant information includes the starting offset of the data block in the target file, the size of the data block, the identifier of the data file space to which the data block belongs, and the offset of the data block in the data file space to which the data block belongs.

[0081] In an implementation of the embodiment of the present invention, the terminal may further merge the first node and the second node (the second node is used to represent the second data block) to generate a target node, the target node is used to represent the target data block. The target node includes the interval of the target data block.

[0082] In an optional implementation, the above Figure 1 The server 102 shown in the figure can also be understood as a storage management server, which is used to store the at least one data file space. The terminal can send a data block insertion request to the storage management server, and the data block insertion request includes the data block to be inserted, and the data block insertion request is used to request the storage management server to insert (or write) the data block to be inserted in the corresponding data file space.

[0083] Optionally, the storage management server may include an upload (put) interface, a download (get) interface and a delete (delete) interface. The upload interface is used to upload data file space, the download interface is used to download data file space, and the delete interface is used to delete data file space.

[0084] For example, Figure 7 The figure is an example of an interval tree provided by an embodiment of the present invention.

[0085] The interval tree may include 10 nodes, namely, node 201 , node 202 , node 203 , node 204 , node 205 , node 206 , node 207 , node 208 , node 209 and node 210 .

[0086] Specifically, the intervals of the data blocks represented by the 10 nodes respectively include [17,22), [9,10), [26,31), [6,9), [16,24), [18,20), [27,29), [1,4), [7,11) and [20,21) respectively.

[0087] Combination Figure 2 ,like Figure 8 As shown, the data storage method provided by the embodiment of the present invention may further include S108-S109.

[0088] S108: The terminal obtains a file identifier, a preset offset, and a preset size of the target file.

[0089] S109: The terminal obtains a target data block set based on the file identifier, the preset start offset, and the preset size of the target file.

[0090] Among them, the target data block set includes M data blocks, the M data blocks belong to the target file, the starting offset of the first data block of the M data blocks in the target file is the same as the preset starting offset, and the ending offset of the last data block of the M data blocks in the target file is the same as a third value, and the third value is the sum of the preset starting offset and the preset size, M≥1.

[0091] It should be understood that the terminal can determine and obtain the target file based on the file identifier of the target file, and then the terminal can also determine the end position of the target data block set (i.e., the third value mentioned above) based on the preset offset and the preset size, and the terminal can also determine the preset start offset as the start position of the target data block set. At this point, the terminal can obtain the target data block set, specifically the M data blocks mentioned above, which are the data blocks included in the target file. The acquisition efficiency and effectiveness of the data blocks can be guaranteed.

[0092] In one implementation of the embodiment of the present invention, the metadata management server may also store data blocks included in the target file. At this time, the terminal may send a data block acquisition request to the metadata management server, and the data block acquisition request includes the file identifier of the target file, the preset start offset, and the preset size. After receiving the data block acquisition request, the metadata management server may determine the M data blocks from all the data blocks included in the target file based on the file identifier of the target file, the preset start offset, and the preset size, and send a data block acquisition response to the terminal. The data block acquisition response includes the target data block set.

[0093] The embodiment of the present invention can divide the functional modules of the terminal and the server according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0094] In the case of dividing each functional module into corresponding functional modules, Fig. 9 A possible structural diagram of the data storage device involved in the above embodiment is shown. Fig. 9 As shown, the data storage device 30 may include: an acquisition module 301 , a processing module 302 , and a deletion module 303 .

[0095] The acquisition module 301 is used to acquire metadata of a target file, the metadata including a start offset of each data block in the target file and a size of each data block, the at least one data block being a data block included in the target file.

[0096] The processing module 302 is used to generate an interval of each data block based on the starting offset of each data block in the target file and the size of each data block, wherein the minimum value in the interval of a data block is the starting offset of the data block in the target file, the maximum value in the interval of the data block is the ending offset of the data block in the target file, the starting offset of the data block in the target file is the distance between the starting position of the data block and the starting position of the target file, and the ending offset of the data block in the target file is the distance between the ending position of the data block and the starting position of the target file.

[0097] The deletion module 303 is used to delete the first overlay data block when the starting offset of the data block to be inserted in the target file is within the interval of the first data block, the first data block is one of the at least one data block, the starting offset of the first overlay data block in the target file is the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file is the ending offset of the first data block in the target file.

[0098] Optionally, the metadata further includes an identifier of the data file space to which each data block belongs and an offset of each data block in the data file space to which each data block belongs; the data storage device 30 further includes a determination module 304 .

[0099] The determination module 304 is used to determine whether the starting offset of the first data block in the target file is the same as the first value, and whether the offset of the first data block in the first data file space is the same as the second value, the first value is the sum of the starting offset of the second data block in the target file and the size of the second data block, the second value is the sum of the offset of the second data block in the first data file space and the size of the second data block, the first data file space is the data file space to which the first data block and the second data block belong, and the second data block is the previous data block of the first data block in the at least one data block.

[0100] The processing module 302 is also used to merge the first data block and the second data block to generate a target data block when the starting offset of the first data block in the target file is the same as the first value and the offset of the first data block in the first data file space is the same as the second value, wherein the ending offset of the target data block in the target file is the same as the ending offset of the first data block in the target file and the starting offset of the target data block in the target file is the same as the starting offset of the second data block in the target file.

[0101] Optionally, the processing module 302 is further configured to generate an initial interval tree based on metadata of the target file, wherein the initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block.

[0102] The processing module 302 is also used to insert the data block to be inserted into the initial interval tree to obtain a target interval tree, wherein the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node, the first node is used to represent the first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

[0103] Optionally, the acquisition module 301 is further used to acquire a file identifier, a preset start offset, and a preset size of the target file.

[0104] The acquisition module 301 is also used to acquire a target data block set based on the file identifier of the target file, the preset start offset and the preset size, the target data block set including M data blocks, the M data blocks belonging to the target file, the start offset of the first data block of the M data blocks in the target file is the same as the preset start offset, the end offset of the last data block of the M data blocks in the target file is the same as a third value, the third value is the sum of the preset start offset and the preset size, M≥1.

[0105] In the case of an integrated unit, Fig.10 FIG. 2 shows a possible structural diagram of the data storage device involved in the above embodiment. Fig.10 As shown, the data storage device 40 may include: a processing module 401 and a communication module 402. The processing module 401 may be used to control and manage the actions of the data storage device 40. The communication module 402 may be used to support the communication between the data storage device 40 and other entities. Fig.10As shown, the data storage device 40 may further include a storage module 403 for storing program codes and data of the data storage device 40 .

[0106] The processing module 401 may be a processor or a controller. The communication module 402 may be a transceiver, a transceiver circuit or a communication interface, etc. The storage module 403 may be a memory.

[0107] When the processing module 401 is a processor, the communication module 402 is a transceiver, and the storage module 403 is a memory, the processor, the transceiver, and the memory may be connected via a bus. The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0108] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0109] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0111] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0113] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A data storage method, characterized in that: include: Acquire metadata of a target file, the metadata including a start offset of each data block in at least one data block in the target file and a size of each data block, the at least one data block being a data block included in the target file; Based on the starting offset of each data block in the target file and the size of each data block, an interval of each data block is generated, wherein the minimum value in the interval of a data block is the starting offset of the data block in the target file, the maximum value in the interval of the data block is the ending offset of the data block in the target file, the starting offset of the data block in the target file is the distance between the starting position of the data block and the starting position of the target file, and the ending offset of the data block in the target file is the distance between the ending position of the data block and the starting position of the target file; In a case where the starting offset of the data block to be inserted in the target file is within the interval of the first data block, deleting a first overlay data block, the first data block being one of the at least one data block, the starting offset of the first overlay data block in the target file being the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file being the ending offset of the first data block in the target file; The metadata also includes an identifier of the data file space to which each data block belongs and an offset of each data block in the data file space to which each data block belongs. The method further includes: Determine whether the start offset of the first data block in the target file is the same as a first value, and whether the offset of the first data block in the first data file space is the same as a second value, the first value is the sum of the start offset of the second data block in the target file and the size of the second data block, the second value is the sum of the offset of the second data block in the first data file space and the size of the second data block, the first data file space is the data file space to which the first data block and the second data block belong, and the second data block is a data block previous to the first data block in the at least one data block; When the starting offset of the first data block in the target file is the same as the first value, and the offset of the first data block in the first data file space is the same as the second value, the first data block and the second data block are merged to generate a target data block, the ending offset of the target data block in the target file is the same as the ending offset of the first data block in the target file, and the starting offset of the target data block in the target file is the same as the starting offset of the second data block in the target file.

2. The data storage method according to claim 1, characterized in that: The method further comprises: Generate an initial interval tree based on the metadata of the target file, wherein the initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block; Insert the data block to be inserted into the initial interval tree to obtain a target interval tree, wherein the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node, the first node is used to represent the first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

3. The data storage method according to claim 1 or 2, characterized in that: The method further comprises: Obtaining a file identifier, a preset start offset, and a preset size of the target file; Based on the file identifier of the target file, the preset start offset and the preset size, a target data block set is obtained, the target data block set includes M data blocks, the M data blocks belong to the target file, the start offset of the first data block of the M data blocks in the target file is the same as the preset start offset, the end offset of the last data block of the M data blocks in the target file is the same as a third value, the third value is the sum of the preset start offset and the preset size, and M≥1.

4. A data storage device, characterized in that: include: Get modules, process modules, and delete modules; The acquisition module is used to acquire metadata of the target file, wherein the metadata includes a start offset of each data block in at least one data block in the target file and a size of each data block, and the at least one data block is a data block included in the target file; The processing module is used to generate an interval of each data block based on the starting offset of each data block in the target file and the size of each data block, wherein the minimum value in the interval of a data block is the starting offset of the data block in the target file, the maximum value in the interval of the data block is the ending offset of the data block in the target file, the starting offset of the data block in the target file is the distance between the starting position of the data block and the starting position of the target file, and the ending offset of the data block in the target file is the distance between the ending position of the data block and the starting position of the target file; The deletion module is configured to delete a first overlay data block when the starting offset of the data block to be inserted in the target file is within the interval of a first data block, the first data block being one of the at least one data block, the starting offset of the first overlay data block in the target file being the starting offset of the data block to be inserted in the target file, and the ending offset of the first overlay data block in the target file being the ending offset of the first data block in the target file; The metadata further includes an identifier of the data file space to which each data block belongs and an offset of each data block in the data file space to which each data block belongs, and the data storage device further includes a determination module; The determination module is used to determine whether the starting offset of the first data block in the target file is the same as a first value, and whether the offset of the first data block in the first data file space is the same as a second value, the first value is the sum of the starting offset of the second data block in the target file and the size of the second data block, the second value is the sum of the offset of the second data block in the first data file space and the size of the second data block, the first data file space is the data file space to which the first data block and the second data block belong, and the second data block is a data block previous to the first data block in the at least one data block; The processing module is further used to merge the first data block and the second data block to generate a target data block when the starting offset of the first data block in the target file is the same as the first value and the offset of the first data block in the first data file space is the same as the second value, wherein the ending offset of the target data block in the target file is the same as the ending offset of the first data block in the target file, and the starting offset of the target data block in the target file is the same as the starting offset of the second data block in the target file.

5. The data storage device according to claim 4, characterized in that: The processing module is further used to generate an initial interval tree based on the metadata of the target file, wherein the initial interval tree includes at least one node, and the at least one node is used to represent the at least one data block; The processing module is also used to insert the data block to be inserted into the initial interval tree to obtain a target interval tree, wherein the first node included in the target interval tree includes the interval of the first uncovered data block, the first node is one of the at least one node, the first node is used to represent the first data block, the starting offset of the first uncovered data block in the target file is the starting offset of the first data block in the target file, and the ending offset of the first uncovered data block in the target file is the starting offset of the data block to be inserted in the target file.

6. The data storage device according to claim 4 or 5, characterized in that: The acquisition module is further used to acquire the file identifier, the preset start offset and the preset size of the target file; The acquisition module is further used to acquire a target data block set based on the file identifier of the target file, the preset start offset and the preset size, the target data block set including M data blocks, the M data blocks belonging to the target file, the start offset of the first data block of the M data blocks in the target file being the same as the preset start offset, the end offset of the last data block of the M data blocks in the target file being the same as a third value, the third value being the sum of the preset start offset and the preset size, and M≥1.

7. An electronic device, characterized in that: The electronic device comprises: processor; a memory configured to store instructions executable by the processor; The processor is configured to execute the instructions to implement the data storage method according to any one of claims 1 to 3.

8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions in the computer-readable storage medium are executed by an electronic device, the electronic device is enabled to execute the data storage method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and readable storage medium

    CN114691681A

  • Data storage method and device, equipment and storage medium

    CN114911410A