A data writing method and a data reading method

By receiving and sorting the data blocks written by users in the user file storage system and sorting them in the logical address order, the problem of read-out resource occupation caused by writing is solved, and the system response speed and data block reading efficiency are improved.

CN114816240BActive Publication Date: 2025-05-06ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210333136.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-05-06
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

When existing user file storage systems write data, file pre-reading requires a lot of resources, affecting the system response speed.

Method used

By receiving the data blocks written by the user, it is stored in memory, and the data blocks are sorted in memory, so that the data blocks of the same user file are continuously and sorted in logical address order, and finally the sorted data is stored persistently.

Benefits of technology

When reading the target data block, the pre-reading process is completed by the way, which reduces the consumption of pre-reading resources, improves the data block reading efficiency, and reduces the system pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816240B_ABST
    Figure CN114816240B_ABST
Patent Text Reader

Abstract

This specification provides a data writing method and a data reading method, receiving a data block written by a user, and storing the data block written by the user in a memory; sorting the data blocks written by the user in the memory to obtain sorted data; wherein, in the sorted data, different data blocks of the same user file are continuous and sorted in the order of logical addresses in the user file; and storing the sorted data persistently. In this way, when the user needs to read a target data block, the target data block and the pre-read data block group are read; the target data block and the pre-read data block group are read; the target data block is the data block to be read, and the pre-read data block group includes at least one data block that is continuous with the physical address of the target data block; the target data block is returned to the user, and the pre-read data block group is stored in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of computer application technology, and in particular, to a data writing method and a data reading method. Background Art

[0002] For a user file storage system, in some cases, a user may write a data block to a file stored by the user in the system (the file stored by the user in the user file storage system is hereinafter referred to as the user file) at regular intervals, and the nodes used to store data generally store the data in the order in which it is received, so that data with consecutive logical addresses in the user file are stored at locations with discontinuous physical addresses in the user file storage system.

[0003] In order to improve the speed at which the user file storage system responds to the user's request to read data, it generally performs pre-reading, that is, when the user reads data at position A of a user file quickly, the data block near position A is read in advance in the memory, so that the user can respond quickly when reading data near position A. However, the above user file writing method requires file pre-reading to occupy more resources. Summary of the invention

[0004] In view of this, one or more embodiments of the present specification provide a data writing method and a data reading method.

[0005] According to a first aspect of one or more embodiments of this specification, a data writing method is provided, the method comprising:

[0006] Receive a data block written by a user, and store the data block written by the user in a memory;

[0007] Arranging the data blocks written by the user in the memory to obtain arranged data; wherein, in the arranged data, different data blocks of the same user file are continuous and arranged in the order of logical addresses in the user file;

[0008] The collated data is persistently stored.

[0009] According to a second aspect of one or more embodiments of the present specification, a data reading method is provided, for reading a desired data block based on the sorted data written by the above data writing method, the method comprising:

[0010] Reading a target data block and a pre-read data block group; the target data block is a data block to be read, and the pre-read data block group includes at least one data block whose physical address is continuous with that of the target data block;

[0011] The target data block is returned to the user, and the pre-read data block group is stored in the memory.

[0012] According to a third aspect of an embodiment of this specification, a data writing device is provided, the device comprising:

[0013] A data block receiving module, used for receiving a data block written by a user, and storing the data block written by the user in a memory;

[0014] A data block sorting module, used for sorting the data blocks written by the user in the memory to obtain sorted data; wherein, in the sorted data, different data blocks of the same user file are continuous and sorted in the order of logical addresses in the user file;

[0015] A persistent storage module is used to persistently store the sorted data.

[0016] According to a fourth aspect of an embodiment of this specification, a data reading device is provided, which is used to read a required data block based on the sorted data written by the above data writing method, and the device includes:

[0017] A reading module, used for reading a target data block and a pre-read data block group; the target data block is a data block to be read, and the pre-read data block group includes at least one data block whose physical address is continuous with that of the target data block;

[0018] The return module is used to return the target data block to the user and store the pre-read data block group in the memory.

[0019] According to a fifth aspect of the embodiments of this specification, a computer device is provided, the computer device comprising:

[0020] one or more processors;

[0021] A memory for storing one or more programs;

[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned data writing method or data reading method.

[0023] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed by a processor, the above-mentioned data writing method or data reading method is implemented.

[0024] According to a seventh aspect of the embodiments of this specification, a computer program is provided, which implements the above-mentioned data writing method or data reading method when executed by a processor.

[0025] This specification provides a data writing method and a data reading method, receiving a data block written by a user, and storing the data block written by the user in a memory; sorting the data blocks written by the user in the memory to obtain sorted data; wherein, in the sorted data, different data blocks of the same user file are continuous and sorted in the order of logical addresses in the user file; and storing the sorted data persistently. In this way, when the user needs to read a target data block, the target data block and the pre-read data block group are read; the target data block and the pre-read data block group are read; the target data block is the data block to be read, and the pre-read data block group includes at least one data block that is continuous with the physical address of the target data block; the target data block is returned to the user, and the pre-read data block group is stored in the memory.

[0026] In this way, the pre-reading process is completed while reading the target data block, and the small block read is converted into a large block read. The pre-reading is completed with very few resources, which reduces the resource consumption caused by the pre-reading and effectively reduces the pressure on the system caused by the data block reading (because if the pre-read data block is used, there is no need to read it again, so the multiple reading process is completed with less resource consumption).

[0027] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.

[0029] Figure 1 is a flow chart of a data writing method according to an exemplary embodiment of the present specification.

[0030] Figure 2A It is a schematic diagram of obtaining sorted data according to an exemplary embodiment of this specification.

[0031] Figure 2B It is a schematic diagram of another method of obtaining sorted data according to an exemplary embodiment of the present specification.

[0032] Figure 3 is a flow chart of a data reading method according to an exemplary embodiment of the present specification.

[0033] Figure 4 It is a schematic diagram of a data writing method and a data reading method according to a specific embodiment of this specification.

[0034] Figure 5 It is a block diagram of a data writing device according to an exemplary embodiment of the present specification.

[0035] Figure 6 is a block diagram of a data reading device according to an exemplary embodiment of the present specification.

[0036] Figure 7 It is a hardware structure diagram of an electronic device in which a data writing device or a data reading device is located according to an exemplary embodiment of this specification. DETAILED DESCRIPTION

[0037] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0038] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0039] The user file storage system is a system that provides cloud file storage services for users. It generally includes a meta node and several data nodes. The meta node is a centralized metadata storage node in a distributed storage system, usually used to store file status information, data block (chunk) location information, data block length information, etc.; the data node is used to store data uploaded by users in the user file storage system.

[0040] In the user file storage system, files uploaded by users are generally called user files (inodes). Each time a user writes, he may write an entire user file or most of the contents of a user file, or he may write a smaller data block each time (for example, a user stores a user file that records the number of web page visits in the user file storage system, and the user needs to obtain the number of web page visits at regular intervals and record the number of web page visits in the user file. Another example is that the user needs to modify data at a certain location, which may also be the case as described above).

[0041] As for the writing method of user files, the writing method is often determined by the size of the data block uploaded by the user at one time. For larger data blocks (i.e. the former mentioned above), in order to facilitate writing, the data block is generally written directly to the data node after the user uploads the data block. After the data node is written, the storage location and other information of the data block are informed to the meta node, so that the meta node can feedback the writing completion to the user. For smaller data blocks, if the same writing method as the larger data blocks is used, since the meta node needs to return to the user, then this method will make it slower to return the writing success to the user. Therefore, in the related technology, for smaller data blocks, the data is generally written to the memory of the meta node first (at this time, the writing success can be returned to the user), and then the meta node stores it persistently in the data node.

[0042] The writing method for smaller data blocks mentioned above is generally called small block writing. For small block writing, since data blocks are generally written in the order in which they are uploaded, in the context of small block writing, data blocks with continuous logical addresses of the same user file will not be uploaded together, which will cause two small data blocks to be stored in discontinuous locations of the physical addresses of the same data node, or stored on different data nodes. In other words, small block writing will cause data fragmentation.

[0043] For reading system fragments, if too many data fragments need to be read in a short period of time, it is necessary to read the data at different locations of the data node, or to read the data through different nodes. The physical addresses of these data are often stored in scattered locations, which can easily consume a lot of resources and cause system instability.

[0044] The user file storage system generally performs pre-reading, which means that the system backend actively reads and caches the data in advance when the user does not explicitly read a block of data. The data that needs to be read for pre-reading is generally determined based on the data currently read by the user. For example, if the user currently reads a data block with a logical address offset (the offset of the logical address relative to the first byte of the user file) of user file 1 of 800k, the user file storage system will pre-read the data block with a logical address offset (offset) of 700k-1000k of the user file in advance (this is just an example and does not limit the pre-read range).

[0045] In the case of small block writes, if pre-reading is required, a large number of data fragments may need to be read in a short period of time, affecting system stability.

[0046] In order to solve the above problem, first of all, we consider that for data nodes, reading several data blocks with consecutive physical addresses at a time will not consume much more resources than reading one data block at a time when the number of data blocks is not particularly large. Therefore, we can use this feature to read several more data blocks with consecutive physical addresses from the data node while reading the data blocks normally. As long as some of the extra data blocks are used, the consumption is worthwhile.

[0047] In order to further increase the number of used data blocks, it is considered that if the data stored in the data node can be sorted so that the logical addresses of data blocks with continuous physical addresses are as adjacent as possible, the number of used data blocks can be increased, so that a small amount of consumption can bring a lot of benefits. However, for the user file storage system, the data node includes several data files (datafile) storing data blocks. The data file is an append only file. This limitation makes it impossible to change the previously written data after the file is written to the data node, that is, it cannot be overwritten, and the data to be written can only be added to the end of the file.

[0048] For a user file storage system with the above characteristics, in order to improve the pre-reading of multiple read data blocks, before writing the data blocks to the data files of the data nodes (that is, before persistent storage), the data blocks are temporarily stored in the memory (if the response speed is improved, they can be stored in the memory of the meta node), and the data blocks written by the user are sorted in the memory, and the data blocks belonging to the same user file are put together and sorted in the order of logical addresses.

[0049] Therefore, when reading a data block normally, data blocks with adjacent physical addresses to the data block can be read at the same time. These data blocks are more likely to be data blocks with close logical addresses, so that the data read at the same time can become effective pre-read data, thereby reducing the consumption of system resources by pre-reading.

[0050] In other words, the present specification provides a data writing method and a data reading method, receiving a data block written by a user, and storing the data block written by the user in a memory; arranging the data blocks written by the user in the memory to obtain arranged data; wherein, in the arranged data, different data blocks of the same user file are continuous and are sorted in the order of the logical addresses in the user file; and storing the arranged data persistently. In this way, when the user needs to read a target data block, the target data block and the pre-read data block group are read; the target data block and the pre-read data block group are read; the target data block is the data block to be read, and the pre-read data block group includes at least one data block that is continuous with the physical address of the target data block; the target data block is returned to the user, and the pre-read data block group is stored in the memory.

[0051] In this way, the pre-reading process is completed while reading the target data block, and the small block read is converted into a large block read. The pre-reading is completed with very few resources, which reduces the resource consumption caused by the pre-reading and effectively reduces the pressure on the system caused by the data block reading (because if the pre-read data block is used, there is no need to read it again, so the multiple reading process is completed with less resource consumption).

[0052] Next, a data writing method shown in this specification will be described in detail.

[0053] like Figure 1 As shown, Figure 1 1 is a flow chart of a data writing method according to an exemplary embodiment of the present specification, comprising the following steps:

[0054] Step 101: receiving a data block written by a user, and storing the data block written by the user in a memory.

[0055] The data block written by the user is the data block written by the user through the client.

[0056] The reason why the data blocks written by the user need to be stored in the memory is that in the data nodes of some cloud file storage systems, the file used to store data blocks is an append only file (see the above description for the specific meaning). Therefore, it is necessary to complete the data sorting before persistent writing. In this case, the data blocks written by the user need to be temporarily stored in the memory to facilitate sorting.

[0057] It should also be noted that the data writing method can be executed by the data node. If the data block is first received by the meta node in order to improve the response speed of the user's write request, then the method can also be executed by the meta node.

[0058] Step 103: Arrange the data blocks written by the user in the memory to obtain arranged data.

[0059] Among them, in the sorted data, different data blocks of the same user file are continuous and sorted according to the order of logical addresses in the user file.

[0060] This step is to organize the data blocks written by the user in the memory so that data blocks with adjacent or close logical addresses can be stored together. In this way, in the process of reading the target data block, the data blocks adjacent to the physical address of the target data block can be read at the same time, and the probability of valid reading of the data blocks read more can be guaranteed to be higher.

[0061] After explaining the overall purpose of step 103, the specific implementation of step 103 will be explained next. Figure 2A As shown, Figure 2A A schematic diagram showing the data written by the user and the resulting collated data, from Figure 2A It can be seen that in the data sorting, the data of the same user file is put together and sorted in the order of logical address (represented by logical address offset) from small to large.

[0062] Next, the execution timing of step 103 will be described. Step 103 can be executed when the number of data blocks written by the user in step 101 exceeds a preset data block number threshold, or when the remaining space in the memory is less than a preset space threshold, or every cycle. This specification does not limit the execution timing of step 103.

[0063] Next, the terms involved in step 103 will be explained. The sorted data is a piece of data obtained by sorting. The data block is the same as the data written into the memory in step 101, but the position of the data block has changed. The meanings of user files and logical addresses are described above, and will not be repeated here. It should be noted that the logical address is generally represented by a logical address offset.

[0064] In addition, in some cases, the cloud file storage system is limited so that each data file used to store user data in the cloud file storage system only stores the data of one user. Figure 2B As shown, in the data sorting, not only the data blocks of the same user file will be continuous, but also the data blocks of different user files of the same user will be continuous, that is, Figure 2B In the example, the data of user file 1 and user file 2 are stored continuously, which makes it easy to persist the sorted data.

[0065] It is also necessary to explain the difference between data files and user files. As mentioned above, user files are logical files, while data files are physical files stored in data nodes and are used to store user data blocks. A user file may be stored in multiple data files, a user file may be stored on multiple data nodes, and a data file can only be stored on one data node.

[0066] Step 105: persistently store the sorted data.

[0067] In step 105, after the arrangement is completed, the data needs to be persistently stored to facilitate data reading. Persistent storage means writing the arranged data into a non-volatile storage medium.

[0068] If the method is executed by a data node, the persistent storage process may be that the data node persistently stores the sorted data in a local disk. If the method is executed by a meta node, the persistent storage process may be that the meta node sends the data to be persisted to a designated data node so that the data node persistently stores the data.

[0069] In the latter case, the data node does not need to change the data writing process, making the implementation of the method more convenient, and the meta node executing the method can return a successful write to the user when the data block is written into the meta node memory, thereby improving the response speed. In the latter case, in other words, the method is applied to the meta node of the user file storage system, and step 105 includes: persistently storing the sorted data in the data node.

[0070] In addition, in order to improve data reliability, a pre-reading table can be added for the sorted data while persistently storing the sorted data. The pre-reading table records the data block sizes of the logical address offsets of each data block with a similar logical address to the same user file in the sorted data. Therefore, during the data reading process, it can be determined based on the pre-reading table whether pre-reading is required (i.e., whether the data block that is continuous with the physical address of the target data block is adjacent to or close to the target data block in terms of the logical address), thereby ensuring the effectiveness of the pre-reading as much as possible and further saving processing resources.

[0071] In other words, the method further comprises: adding a pre-reading table for the sorted data; wherein the pre-reading table stores the logical addresses and sizes of the data blocks in the order of the data blocks in the sorted data.

[0072] As mentioned above, the logical address can be represented by the logical address offset. Since the user file is a logical file, it only has a stored physical address. The logical address is just a figurative concept used to describe the location of the data block in the user file. It can be represented by the logical address offset. Of course, it can also be represented by other methods. This specification does not limit this. Size refers to the size of the data block.

[0073] The reason why the order of the data blocks in the sorted data is followed is because it is necessary to indicate whether the logical addresses of the data blocks with consecutive physical addresses in the sorted data are close to each other.

[0074] In addition, in order to further reduce the space occupied by the pre-reading table and improve processing efficiency, the logical address and size of the data blocks included in the user file may not be written in the pre-reading table in the following situations: first, a user file corresponds to only one data block in the sorted data; second, a user file corresponds to more data blocks in the sorted data.

[0075] The purpose of doing this is that if the sorted data only includes one data block of a certain user file, then when reading the target data block, it is only necessary to determine whether the logical address of the target data block is recorded in the pre-reading table. If not, pre-reading can be disabled for the target data block to reduce the resources consumed by searching the pre-reading table (if these are not recorded in the pre-reading table, it is also necessary to determine whether the data blocks continuous with the physical address of the target data block need to be pre-read based on the pre-reading table).

[0076] If a user file has many data blocks in the sorted data, the user client may find this phenomenon and perform pre-reading in the upper layer. In order to prevent duplication of work, it is assumed that these data blocks are not pre-read. In this way, if the logical address of the target data block is not recorded in the pre-reading table, it is directly determined that pre-reading is not enabled for the target data block.

[0077] In other words, when the number of data blocks of any user file included in the sorted data exceeds a preset threshold, or when the number of data blocks of any user file included in the sorted data is less than 2, the pre-reading table does not include the logical address and size of the data block of the user file.

[0078] Finally, the storage location of the pre-reading table needs to be explained. The pre-reading table can be added to the header of the data file, or the pre-reading tables of all data files can be stored in the data node. However, when the above method is executed by the meta node, the pre-reading table is generally not stored in the meta node. The specific reasons are as follows. First, there is a lot of data in the pre-reading table. If it is all stored in the meta node, it will take up a lot of storage space in the meta node. Second, the pre-reading table and other data stored in the meta node are not at the same level. This will easily cause data at multiple levels to be mixed together, disrupting the hierarchical logic of the system. Third, in some cases, placing the pre-reading table in the meta node will incur a large overhead. For example, when merging multiple data files, it is also necessary to merge the pre-reading tables corresponding to the data files. Then, merging the data files will require the meta node to execute the process of merging the pre-reading tables, which will bring greater processing pressure to the meta node.

[0079] After describing the data writing method, a data reading method shown in an embodiment of this specification will be described next. This method reads the required data block based on the sorted data written by the above-mentioned data writing method.

[0080] like Figure 3 As shown, Figure 3 is a schematic diagram of a data reading method according to an exemplary embodiment of the present specification, comprising the following steps:

[0081] Step 301, read the target data block and the pre-read data block group.

[0082] The target data block is a data block to be read, and the pre-read data block group includes at least one data block whose physical address is continuous with that of the target data block.

[0083] In other words, in step 301, when reading the target data block, the small block read is converted into a large block read, and the pre-read data block group is read at the same time, thereby saving the resources required for reading through small resource consumption.

[0084] In the above data writing method, as long as the data blocks belong to the same user file, they will be put together in the sorted data, so the data blocks whose physical addresses are continuous with the physical addresses of the target data blocks are likely to be the data blocks that may be read later. On the basis of increasing consumption only slightly, greater resource consumption can be reduced.

[0085] Among them, the number of data blocks included in the pre-read data block group and the relative physical position to the target data block can be pre-defined. For example, the pre-defined pre-read data block group includes 2 data blocks, which are the data blocks before and after the physical address of the target data block. For another example, the pre-defined pre-read data block group includes 2 data blocks, which are the two data blocks after the physical address of the target data block.

[0086] In addition, the number of data blocks included in the pre-read data block group can be determined according to the remaining space in the memory. The more remaining space there is, the more data blocks there are in the pre-read data block group. In other words, step 301 includes: determining the number of data blocks in the pre-read data block group according to the current remaining space in the memory, and reading the target data block and the pre-read data block group.

[0087] Considering that when a user reads a data block with a logical address A, the possibility of continuing to read the data after A is greater than the possibility of reading the data before A, the data blocks included in the pre-read data block group can also be determined as data blocks after the physical address of the target data block. In other words, step 301 includes: reading the target data block and the pre-read data block group, the logical address of the data block included in the pre-read data block group in the user file is greater than the logical address of the target data block in the user file.

[0088] In addition, it is also possible to determine how many times the pre-read data block group has been used by the user according to the historical reading record of the user. If it is used more often, it can be determined to enable pre-reading and pre-read more data blocks. If it is used less often, the scope of pre-reading can be reduced or even pre-reading is not enabled for the target data block. In other words, step 301 includes: determining the number of data blocks of the pre-read data block group according to the usage of other pre-read data block groups in the historical reading record of the user, and reading the target data block and the pre-read data block group.

[0089] Furthermore, when a pre-reading table is stored, it is also possible to determine whether to enable pre-reading for the target data block each time the target data block is read, and to determine the number and positions of the data blocks included in the pre-reading data block group when pre-reading is enabled.

[0090] Specifically, the pre-reading range can be determined according to the pre-reading table. For example, if two data blocks with consecutive physical addresses in the pre-reading table belong to the same user file, but the logical addresses of the two data blocks are far apart, then the data block will not be pre-read. If a user file has many data blocks with consecutive physical addresses in the pre-reading table, which data blocks to read can be determined according to the logical addresses of each data block. For example, the read-only logical address offset can be limited to data blocks within 1000k relative to the target data block.

[0091] In other words, a pre-read table is also stored, in which the logical addresses and sizes of the data blocks are stored according to the order of the data blocks in the sorted data. Step 301 includes: determining the data blocks included in the pre-read data block group according to the logical addresses of other data blocks that are continuous with the physical address of the target data block recorded in the pre-read table; and reading the target data block and the pre-read data block group.

[0092] Step 303: Return the target data block to the user, and store the pre-read data block group in the memory.

[0093] In this step, the target data block is the data block required by the user, so it needs to be returned to the user. Since the speed of accessing the non-volatile storage medium is slower than the speed of accessing the memory, pre-reading can be achieved by storing the pre-read data block group in the memory, so that when the user needs to read the pre-read data later, the corresponding data can be quickly returned to the user.

[0094] In addition, it should be noted that since the pre-read data block group is stored in the memory, and the memory space is limited, the space occupied by the pre-read data block group cannot be expanded indefinitely. When the space occupied by the pre-read data block group is greater than a certain value, the pre-read data block group with the earliest write time can be deleted; or one or more pre-read data block groups with an earlier write time can be deleted at regular intervals to prevent excessive pre-reading from affecting the normal operation of the system.

[0095] Finally, the execution subject of the above method will also be described. Similar to the data writing method, the data reading method can be executed by a data node or a meta node. It should be noted that when executed by a data node, the data node can be directly connected to the user, and step 303 means that the data node directly returns the target data to the user. When executed by a meta node, step 301 specifically obtains the target data block and the pre-read data block group from the data node, and if the pre-read data block group is determined based on the pre-read table, then the data node is used to determine which data blocks are included in the pre-read data block group through the pre-read table. And step 303 is stored in the memory of the meta node.

[0096] It should also be noted that the size of the target data block must be smaller than or equal to the size of the data block written by the user.

[0097] Next, the data writing method and the data reading method shown in this specification will be described through a specific embodiment.

[0098] like Figure 4 As shown, Figure 4 It is a schematic diagram of a data writing method and a data reading method according to a specific embodiment of this specification.

[0099] The user first sends the data blocks to be written to the meta node in small block write mode. After receiving these data blocks, the meta node stores them in the local memory. Later, if persistent storage is required, the data blocks written by the user are first sorted to obtain sorted data. The sorted data is sorted in the order of user files and logical address offsets of the data blocks in the user files. After the sorting is completed, the sorted data is persistently written to the data file of the specified data node.

[0100] At the same time, a pre-read table is added to the specified data file. The array table contains an array. Each item in the array is a range of data blocks with continuous physical addresses of a user file on the data file. For example, it can be in the following form:

[0101]

[0102] That is, it means that there are 5 data blocks with consecutive physical addresses in the data file, namely, a 4K data block at offset = 800k of user file 1, a 4K data block at offset = 900k of user file 1, and a 4K data block at offset = 1000k of user file 1.

[0103] If a user file has a single data block in the sorted data (i.e., other data blocks in the sorted file that are physically continuous with the single data block and the single data block do not belong to the same user file), the logical address offset and size of the single data block will not be recorded. If a user file has multiple data blocks with consecutive physical addresses in the sorted data, and the number of these data blocks exceeds the preset number threshold, the information of these data blocks will not be recorded in the pre-read table.

[0104] After explaining the data writing, the data reading process will be explained next.

[0105] When the meta node receives a data read request from a user, if the target data block needs to be read in small blocks, the data node will determine whether the target data block is in the pre-read table. If not, only the target data block will be read. If it is, it can determine whether to read these data blocks based on the logical address offsets of the data blocks in the pre-read table that are continuous with the physical address of the target data block (the other data blocks read except the target data block are collectively referred to as the pre-read data block group). For example, continuing with the above example, the target data block to be read is a 2k data block at the offset = 800k position of user file 1. Then the several data blocks that are continuous with the physical address of the target data block and the logical address offset of the target data block are not much different. Then the small block read can be converted into a large block read, and three data blocks can be read together and sent to the meta node.

[0106] The meta node returns the target data block to the user and stores other data blocks in its own memory for future use.

[0107] Corresponding to the embodiments of the aforementioned method, this specification also provides embodiments of a device and a terminal to which it is applied.

[0108] like Figure 5 As shown, Figure 5is a block diagram of a data writing device according to an exemplary embodiment of the present specification, the device comprising:

[0109] The data block receiving module 510 is used to receive the data block written by the user and store the data block written by the user into the memory.

[0110] The data block sorting module 520 is used to sort the data blocks written by the user in the memory to obtain sorted data; wherein, in the sorted data, different data blocks of the same user file are continuous and sorted according to the order of logical addresses in the user file.

[0111] The persistent storage module 530 is used to persistently store the sorted data.

[0112] In an optional embodiment, the device also includes: a pre-reading table adding module 540 (not shown in the figure), which is used to add a pre-reading table for the sorted data; wherein the pre-reading table stores the logical address and size of the data block in the order of each data block in the sorted data.

[0113] In an optional embodiment, when the number of data blocks of any user file included in the sorted data exceeds a preset threshold value, or when the number of data blocks of any user file included in the sorted data is less than 2, the pre-reading table does not include the logical address and size of the data block of the user file.

[0114] In an optional embodiment, the method is applied to a meta node of a user file storage system; the persistent storage module 530 is used to: persistently store the sorted data in a data node.

[0115] like Figure 6 As shown, Figure 6 1 is a block diagram of a data reading device according to an exemplary embodiment of the present specification, the device is used to read the required data block based on the sorted data written by the aforementioned data writing method, and the device includes:

[0116] The reading module 610 is used to read the target data block and the pre-read data block group; the target data block is the data block to be read, and the pre-read data block group includes at least one data block that is continuous with the physical address of the target data block.

[0117] The return module 620 is used to return the target data block to the user and store the pre-read data block group in the memory.

[0118] In an optional embodiment, a pre-read table is also stored, in which the logical addresses and sizes of the data blocks are stored in the order of the data blocks in the sorted data. The reading module 610 is used to determine the data blocks included in the pre-read data block group according to the logical addresses of other data blocks that are continuous with the physical address of the target data block recorded in the pre-read table; and read the target data block and the pre-read data block group.

[0119] In an optional embodiment, the reading module 610 is used to determine the number of data blocks in the pre-read data block group based on the current remaining memory space, and read the target data block and the pre-read data block group; or, read the target data block and the pre-read data block group, the logical address of the data block included in the pre-read data block group in the user file is larger than the logical address of the target data block in the user file; or, determine the number of data blocks in the pre-read data block group based on the usage of other pre-read data block groups in the user's historical reading records, and read the target data block and the pre-read data block group.

[0120] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0121] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying creative labor.

[0122] like Figure 7 As shown, Figure 7 A hardware structure diagram of a computer device in which a data writing device or a data reading device is located is shown, and the device may include: one or more processors 1010, a memory 1020 for storing one or more programs, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0123] The processor 1010 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits. When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned data writing method or data reading method.

[0124] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0125] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0126] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0127] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0128] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0129] The embodiments of the present specification also provide a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned data writing method or data reading method is implemented.

[0130] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0131] The present specification also provides a computer program, which implements the above-mentioned data writing method or data reading method when executed by a processor.

[0132] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0133] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A data writing method, the method comprising: Receive a data block written by a user, and store the data block written by the user in the memory of a metanode of the cloud storage system; Arranging the data blocks written by the user in the memory to obtain arranged data; wherein, in the arranged data, different data blocks of the same user file are continuous and arranged in the order of logical addresses in the user file; Persistently storing the sorted data in a data node; A pre-reading table is added for the sorted data; wherein the pre-reading table stores the logical address and size of the data block according to the order of each data block in the sorted data, and the pre-reading table is used to determine the pre-read data block group according to the pre-reading table when reading the target data block; When the number of data blocks of any user file included in the sorted data exceeds a preset threshold, or when the number of data blocks of any user file included in the sorted data is less than 2, the pre-reading table does not include the logical address and size of the data block of the user file.

2. A data reading method, for reading a desired data block based on the sorted data written by the data writing method according to claim 1, the method comprising: Read the target data block and the pre-read data block group; The target data block is a data block to be read, and the pre-read data block group includes at least one data block whose physical address is continuous with the target data block; The target data block is returned to the user, and the pre-read data block group is stored in the memory.

3. The method according to claim 2, further comprising storing a pre-reading table, wherein the pre-reading table stores the logical addresses and sizes of the data blocks according to the order of the data blocks in the sorted data; The reading of the target data block and the pre-reading data block group includes: Determine the data blocks included in the pre-read data block group according to the logical addresses of other data blocks recorded in the pre-read table that are continuous with the physical address of the target data block; Read the target data block and the pre-read data block group.

4. The method according to claim 2, wherein the reading of the target data block and the pre-reading of the data block group comprises: According to the current remaining memory space, determine the number of data blocks of the pre-read data block group, and read the target data block and the pre-read data block group; Or, reading the target data block and the pre-read data block group, wherein the logical address of the data block included in the pre-read data block group in the user file is greater than the logical address of the target data block in the user file; Or, according to the usage of other pre-read data block groups in the user's historical read records, the number of data blocks in the pre-read data block group is determined, and the target data block and the pre-read data block group are read.

5. A data writing device, comprising: A data block receiving module, used for receiving data blocks written by users, and storing the data blocks written by users in the memory of a metanode of a cloud storage system; A data block sorting module, used for sorting the data blocks written by the user in the memory to obtain sorted data; wherein, in the sorted data, different data blocks of the same user file are continuous and sorted according to the order of logical addresses in the user file; A persistent storage module, used for persistently storing the sorted data in a data node; A pre-reading table adding module, used for adding a pre-reading table for the sorted data; wherein the pre-reading table stores the logical address and size of the data block according to the order of each data block in the sorted data; When the number of data blocks of any user file included in the sorted data exceeds a preset threshold, or when the number of data blocks of any user file included in the sorted data is less than 2, the pre-reading table does not include the logical address and size of the data block of the user file.

6. A data reading device, for reading a desired data block based on the sorted data written by the data writing method according to claim 1, the device comprising: A reading module, used for reading a target data block and a pre-read data block group; The target data block is a data block to be read, and the pre-read data block group includes at least one data block whose physical address is continuous with the target data block; The return module is used to return the target data block to the user and store the pre-read data block group in the memory.

7. A computer device, comprising: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

8. A computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 4.

9. A computer program product, wherein when the computer program product is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Dynamic granule-based intermediate storage

    CN104049908A