Data management method and device
By optimizing the distribution and storage method of data files in storage blocks, the problem of space waste and performance degradation caused by direct storage in massive small file application scenarios is solved, and more efficient data storage management is achieved.
Patent Information
- Application Number
- CN202510015931.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
In the application scenarios of massive small files, the direct storage method leads to waste of space, performance and efficiency reduction. How to optimize data storage management has become an urgent problem.
By responding to data writing instructions, the thread number information and data length information of the storage block are determined, the file offset is determined according to the data writing order and the memory size of the storage block, and the distribution and storage method of data files in the storage block are optimized.
It realizes efficient organization and management of data files, improves the utilization rate of storage space, and reduces storage costs and performance bottlenecks.
Smart Images

Figure CN119988335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a data management method and device. Background Art
[0002] In the modern computing environment, data transmission management is a critical requirement that runs through various application scenarios, whether it is daily use of personal computers or data processing in large data centers. However, this data management process is not simple and direct.
[0003] In the prior art, direct storage is usually used to manage data. In the case of massive small files, direct storage will cause a lot of space waste, resulting in insufficient space utilization. This is mainly because small files themselves do not occupy much space, but the storage system needs to allocate certain metadata and management information to each file, resulting in reduced storage system performance and efficiency.
[0004] Therefore, how to manage and optimize data storage is a technical problem that current technicians urgently need to solve. Summary of the invention
[0005] The present invention aims to provide a data management method and device to solve the technical problem of how to manage and optimize data storage.
[0006] In order to solve the above technical problems, an embodiment of the present invention provides a data management method, including:
[0007] In response to a data write instruction, respectively determining thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created by a local memory and has a unique thread number;
[0008] Determine the corresponding directory index information according to the data writing sequence information in the data file to be written; determine the file offset according to the directory index information and the current memory size of the storage block;
[0009] Performing a write operation on the data file to be written according to the data length information, the directory index information and the file offset;
[0010] In response to the data read instruction, a read operation is performed on the data file to be written according to the directory index information, the file offset and the data length information in sequence.
[0011] In a preferred embodiment of the present invention, performing a write operation on the data file to be written according to the data length information, the directory index information and the file offset includes:
[0012] Determine, according to the directory index information, a target storage block corresponding to the data file to be written;
[0013] Determine the starting position of the data file to be written in the target storage block according to the file offset;
[0014] A write operation is performed on the data file to be written based on the starting position.
[0015] In a preferred embodiment of the present invention, before performing a write operation on the to-be-written data file based on the starting position, the method further includes:
[0016] Obtaining memory capacity information of the target storage block;
[0017] The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
[0018] In a preferred embodiment of the present invention, the step of performing a read operation on the data file to be written according to the directory index information, the file offset and the data length information in sequence includes:
[0019] Locating a storage block corresponding to the directory index information in the local memory;
[0020] According to the file offset, obtaining the starting position of the data file to be written in the storage block;
[0021] Based on the length information and the starting position, a read operation is performed on the data file to be written.
[0022] In a preferred embodiment of the present invention, after performing a read operation on the data file to be written, the data management method further includes:
[0023] Record the last modification time of the data file to be written;
[0024] Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data;
[0025] A deletion operation is performed on the expired data.
[0026] Another embodiment of the present invention provides a data management device, including:
[0027] An acquisition module, configured to respectively determine, in response to a data write instruction, thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created from a local memory and has a unique thread number;
[0028] An extraction module, used to determine corresponding directory index information according to the data writing sequence information in the data file to be written; and determine the file offset according to the directory index information and the current memory size of the storage block;
[0029] A writing module, used for performing a writing operation on the data file to be written according to the data length information, the directory index information and the file offset;
[0030] The reading module is used to perform a reading operation on the data file to be written according to the directory index information, the file offset and the data length information in sequence when responding to the data reading instruction.
[0031] In a preferred embodiment of the present invention, the writing module is specifically used for:
[0032] Determine, according to the directory index information, a target storage block corresponding to the data file to be written;
[0033] Determine the starting position of the data file to be written in the target storage block according to the file offset;
[0034] A write operation is performed on the data file to be written based on the starting position.
[0035] In a preferred embodiment of the present invention, the writing module is further used for:
[0036] Obtaining memory capacity information of the target storage block;
[0037] The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
[0038] In a preferred embodiment of the present invention, the reading module is specifically used for:
[0039] Locating a storage block corresponding to the directory index information in the local memory;
[0040] According to the file offset, obtaining the starting position of the data file to be written in the storage block;
[0041] Based on the length information and the starting position, a read operation is performed on the data file to be written.
[0042] In a preferred embodiment of the present invention, a cleaning module is further included, which is used to:
[0043] Record the last modification time of the data file to be written;
[0044] Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data;
[0045] A deletion operation is performed on the expired data.
[0046] Compared with the prior art, the beneficial effects of the present invention are at least one of the following:
[0047] (1) The present invention can efficiently organize and manage the distribution of data files in storage blocks by clarifying thread number information, data length information and data writing sequence information, so that the writing and reading operations of data files are more orderly and efficient.
[0048] (2) The present invention determines the file offset according to the current memory size of the storage block, which can fully utilize the storage space, avoid memory fragmentation problems, improve the utilization rate of the storage space, and reduce storage costs.
[0049] (3) By introducing directory index information, the present invention makes it possible to locate data files more quickly. When reading data, the required data can be quickly found directly based on the directory index information, file offset and data length information, thereby improving data access speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flow chart of a data management method in one embodiment of the present invention;
[0051] Figure 2 It is a structural diagram of a data management device in one embodiment of the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] In the description of this application, the terms "first", "second", "third", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the feature. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0054] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. The terms "vertical", "horizontal", "left", "right", "upper", "lower" and similar expressions used herein are only for illustrative purposes, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0055] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood by specific circumstances.
[0056] An embodiment of the present invention provides a data management method. For details, see Figure 1 , Figure 1 The flowchart of the data management method in one embodiment of the present invention is shown, which includes steps S1-S4:
[0057] S1: In response to a data write instruction, respectively determine thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created by a local memory and has a unique thread number;
[0058] It is understandable that in the process of storing a large amount of small file data of random size, in order to improve the efficiency and response speed of data processing, the data is usually divided and stored in multiple storage blocks, which can be distributed on different nodes or memory areas. Each storage block can be assigned to a specific thread or processing unit for processing, so that the parallel processing capabilities of multi-core processors can be fully utilized. In this scenario, when responding to a data write instruction, it is necessary to determine the thread number information and data length information of the storage block corresponding to the data file to be written.
[0059] Among them, massive: generally refers to a data volume of more than 10 billion; small files of random size: files of random size refer to files of uncertain size, and small files generally refer to files of a few bytes to 1MB. In one embodiment of the present application, the file size can be random, and most files are small files. The size of the storage block is fixed, for example: 1GB, and can be freely configured. When small file data is written in, the data of this small file is written into the generated storage block.
[0060] Specifically, when a data write instruction is received, the instruction usually contains the data file to be written and its content. Each storage block has a unique thread number (or processing unit number), which identifies the thread or processing unit responsible for processing the storage block. It is necessary to determine which thread or processing unit should be responsible for processing a part of the data file to be written (i.e., a storage block) based on a certain strategy (such as load balancing, data locality, etc.).
[0061] Preferably, in one embodiment of the present invention, before performing a write operation on the data file to be written, the method further includes:
[0062] Obtaining memory capacity information of the target storage block;
[0063] The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
[0064] The data length information of the data file to be written refers to the total size of the data file or the amount of data to be written. That is, before writing the small file data into the storage block, the data block may already have a data file, and when the capacity of the storage block has increased to the maximum capacity, a new storage block will be added for data storage.
[0065] S2: determining corresponding directory index information according to the data writing sequence information in the data file to be written; determining a file offset according to the directory index information and the current memory size of the storage block;
[0066] When executing data writing, since the storage device (local memory) is a shared resource, multiple threads operating the same file concurrently will inevitably cause competition. In one embodiment of the present application, competition is handled by adding a layer of directory, and the directory is the corresponding thread number. For example: Assume that the thread number is 1, the data directory is / data / 1 / file.data, and 1 in the path is the thread number. Each thread has a different thread number, thereby avoiding competition for writing data, improving the speed of writing, and preventing data confusion caused by competition.
[0067] The directory index information is used to record the writing order and position of data in the data file. During the data writing process, the directory index is generated or updated based on the writing order information of the data (for example, the number of the data block or record, the timestamp, etc.). The file offset refers to the relative position of the data file in the storage block, and is usually used to locate the data to be read or written. During the data writing process, the file offset is determined based on the directory index information and the current memory size of the storage block. This offset is usually in bytes and is relative to the beginning of the file.
[0068] S3: performing a write operation on the data file to be written according to the data length information, the directory index information and the file offset;
[0069] Preferably, in an embodiment of the present invention, performing a write operation on the to-be-written data file according to the data length information, the directory index information and the file offset comprises:
[0070] Determine, according to the directory index information, a target storage block corresponding to the data file to be written;
[0071] Determine the starting position of the data file to be written in the target storage block according to the file offset;
[0072] A write operation is performed on the data file to be written based on the starting position.
[0073] In this step, the target storage block corresponding to the data file to be written is determined according to the mapping relationship between the data file and the storage block provided by the directory index information. After the target storage block is determined, the starting position of the data file to be written in the target storage block is determined according to the file offset, wherein the file offset is an offset relative to the starting position of the storage block, which indicates the position of the storage block to which the data should be written. At the determined starting position, the write operation is performed based on the actual data length of the data file to be written.
[0074] In a real-time example of the present invention, the write operation can be implemented in a variety of ways, including:
[0075] Direct write: Write the contents of the data file to be written to the starting position of the target storage block. This method is suitable for simple storage systems or scenarios with low requirements for data writing performance.
[0076] Buffered writing: first write the contents of the data file to be written into a buffer, and then write the data in the buffer into the target storage block at one time. This method can reduce the frequent access to the storage block and improve the writing speed.
[0077] Concurrent writing: Split the data file to be written into multiple data blocks, and write these data blocks to different storage nodes concurrently. This method can fully utilize the parallel processing capabilities of the storage system and improve writing efficiency.
[0078] S4: In response to the data read instruction, a read operation is performed on the data file to be written according to the directory index information, the file offset and the data length information in sequence.
[0079] Preferably, in an embodiment of the present invention, performing a read operation on the to-be-written data file according to the directory index information, the file offset and the data length information in sequence includes:
[0080] Locating a storage block corresponding to the directory index information in the local memory;
[0081] According to the file offset, obtaining the starting position of the data file to be written in the storage block;
[0082] Based on the length information and the starting position, a read operation is performed on the data file to be written.
[0083] Specifically, according to the directory index information, the corresponding storage block is quickly found in the local memory. After the storage block is determined, the starting position of the required data segment is found inside the storage block according to the file offset. The length information indicates the length of the data segment to be read, that is, it is used to indicate how much data needs to be read. After determining the starting position and length, the read operation is performed on the data file to be written.
[0084] Preferably, in one embodiment of the present invention, after performing a read operation on the data file to be written, the data management method further includes:
[0085] Record the last modification time of the data file to be written;
[0086] Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data;
[0087] A deletion operation is performed on the expired data.
[0088] In this step, the data retention time is a threshold defined by the system administrator or policy, indicating the maximum time that the data file should be retained in the system. This time can be set according to the nature of the data, business needs or regulatory requirements.
[0089] In order to optimize the use of storage resources and avoid unnecessary data accumulation, it is necessary to implement an effective data lifecycle management strategy. Specifically, after performing a read operation or any other modification operation, record the last modification time of the data file. Scan the data files in the storage regularly or as needed, obtain the last modification time of each file, and compare it with the preset data retention time. If the last modification time of a file is earlier than the data retention time, the system should mark the file as expired data. This means that the file has exceeded its predetermined life cycle and can be safely deleted. For data files marked as expired, perform a delete operation to remove them from storage. This step is key to freeing up storage space, reducing data redundancy, and improving system performance.
[0090] Compared with the prior art, the beneficial effects of the present invention are at least one of the following:
[0091] (1) The present invention can efficiently organize and manage the distribution of data files in storage blocks by clarifying thread number information, data length information and data writing sequence information, so that the writing and reading operations of data files are more orderly and efficient.
[0092] (2) The present invention determines the file offset according to the current memory size of the storage block, which can fully utilize the storage space, avoid memory fragmentation problems, improve the utilization rate of the storage space, and reduce storage costs.
[0093] (3) By introducing directory index information, the present invention makes it possible to locate data files more quickly. When reading data, the required data can be quickly found directly based on the directory index information, file offset and data length information, thereby improving data access speed.
[0094] Another embodiment of the present invention provides a data management device. For details, see Figure 2 , Figure 2The structure diagram of the data management device in one embodiment of the present invention is shown, which includes: an acquisition module 11, an extraction module 12, a writing module 13, and a reading module 14, wherein:
[0095] The acquisition module 11 is used to determine, in response to a data writing instruction, thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created by a local memory and has a unique thread number;
[0096] The extraction module 12 is used to determine the corresponding directory index information according to the data writing order information in the data file to be written; and determine the file offset according to the directory index information and the current memory size of the storage block;
[0097] A writing module 13, configured to perform a writing operation on the data file to be written according to the data length information, the directory index information and the file offset;
[0098] The reading module 14 is used to perform a reading operation on the data file to be written according to the directory index information, the file offset and the data length information in sequence when responding to the data reading instruction.
[0099] Preferably, in one embodiment of the present invention, the writing module is specifically used for:
[0100] Determining a target storage block corresponding to the data file to be written according to the directory retrieval information;
[0101] Determine the starting position of the data file to be written in the target storage block according to the file offset;
[0102] A write operation is performed on the data file to be written based on the starting position.
[0103] Preferably, in one embodiment of the present invention, the writing module is further used for:
[0104] Obtaining memory capacity information of the target storage block;
[0105] The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
[0106] Preferably, in one embodiment of the present invention, the reading module is specifically used for:
[0107] Locating a storage block corresponding to the directory index information in the local memory;
[0108] According to the file offset, obtaining the starting position of the data file to be written in the storage block;
[0109] Based on the length information and the starting position, a read operation is performed on the data file to be written.
[0110] Preferably, in one embodiment of the present invention, a cleaning module is further included, for:
[0111] Record the last modification time of the data file to be written;
[0112] Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data;
[0113] A deletion operation is performed on the expired data.
[0114] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A data management method, characterized in that: include: In response to a data write instruction, respectively determining thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created by a local memory and has a unique thread number; Determine the corresponding directory index information according to the data writing sequence information in the data file to be written; determine the file offset according to the directory index information and the current memory size of the storage block; Performing a write operation on the data file to be written according to the data length information, the directory index information and the file offset; In response to the data read instruction, a read operation is performed on the data file to be written according to the directory index information, the file offset and the data length information in sequence.
2. The data management method according to claim 1, characterized in that: The performing a write operation on the to-be-written data file according to the data length information, the directory index information and the file offset comprises: Determining a target storage block corresponding to the data file to be written according to the directory index information; Determine the starting position of the data file to be written in the target storage block according to the file offset; A write operation is performed on the data file to be written based on the starting position.
3. The data management method according to claim 2, characterized in that: Before performing a write operation on the to-be-written data file based on the starting position, the method further includes: Obtaining memory capacity information of the target storage block; The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
4. The data management method according to claim 1, characterized in that: The step of performing a read operation on the data file to be written according to the directory index information, the file offset and the data length information in sequence includes: Locating a storage block corresponding to the directory index information in the local memory; According to the file offset, obtaining the starting position of the data file to be written in the storage block; Based on the length information and the starting position, a read operation is performed on the data file to be written.
5. The data management method according to claim 1, characterized in that: After performing a read operation on the data file to be written, the data management method further includes: Record the last modification time of the data file to be written; Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data; A deletion operation is performed on the expired data.
6. A data management device, characterized in that: include: An acquisition module, configured to respectively determine, in response to a data write instruction, thread number information of a storage block corresponding to a data file to be written and data length information of the data file to be written, wherein the storage block is created from a local memory and has a unique thread number; An extraction module, used to determine corresponding directory index information according to the data writing sequence information in the data file to be written; and determine the file offset according to the directory index information and the current memory size of the storage block; A writing module, used for performing a writing operation on the data file to be written according to the data length information, the directory index information and the file offset; The reading module is used to perform a reading operation on the data file to be written according to the directory index information, the file offset and the data length information in sequence when responding to the data reading instruction.
7. The data management device according to claim 6, characterized in that: The writing module is specifically used for: Determining a target storage block corresponding to the data file to be written according to the directory index information; Determine the starting position of the data file to be written in the target storage block according to the file offset; A write operation is performed on the data file to be written based on the starting position.
8. The data management device according to claim 7, characterized in that: The writing module is further used for: Obtaining memory capacity information of the target storage block; The data length information is compared with the memory capacity information. If the data length information is greater than the memory capacity information, a storage block other than the target storage block is selected to perform a write operation.
9. The data management device according to claim 7, characterized in that: The reading module is specifically used for: Locating a storage block corresponding to the directory index information in the local memory; According to the file offset, obtaining the starting position of the data file to be written in the storage block; Based on the length information and the starting position, a read operation is performed on the data file to be written.
10. The data management device according to claim 7, characterized in that: Also includes cleanup modules for: Record the last modification time of the data file to be written; Scanning the data file to be written, comparing the last modification time with the preset data retention time, and if the last modification time is earlier than the data retention time, marking the data file to be written as expired data; A deletion operation is performed on the expired data.