Virtual file system suitable for high-speed access of batch small-size files
By designing a virtual file system, a large number of small-sized files are stored in a large virtual file and only searched and sorted in the file index table, the performance degradation and resource consumption problems in the processing of small-sized files in the prior art are solved, and efficient file management and storage are achieved.
Patent Information
- Application Number
- CN202510138640.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
When processing large number of small-sized files, existing data management systems face problems such as performance degradation, disk fragmentation, resource consumption, backup and recovery difficulties, and space waste.
A virtual file system is designed to store a large number of small-sized files in a large virtual file through virtual file body and virtual file reading and writing programs, and only search and sort in the file index table to avoid global retrieval and improve efficiency.
Through the virtual file system, the retrieval and deletion efficiency of a large number of small-sized files is significantly improved, resource consumption is reduced, storage space is optimized, and backup and recovery process is simplified.
Smart Images

Figure CN120067071A_ABST
Abstract
Description
Technical Field
[0001] This solution relates to the field of data storage, and specifically to a virtual file system suitable for high-speed access to a batch of small-sized files. Background Art
[0002] Existing operating systems provide various data storage mechanisms to meet different needs. However, most current data management systems are designed to adapt to large-sized files and are not optimized for a large number of small-sized files. When the data storage system of an operating system processes a large number of small-sized files, the following problems may be encountered:
[0003] Performance degradation: The metadata of the file system (such as directory structure, file attributes, etc.) will increase significantly as the number of files increases, which may lead to slower read and write operations because the system needs to spend more time to search for and update this information.
[0004] Disk fragmentation: Frequent creation and deletion of small-sized files will cause disk fragmentation, which in turn affects the file access speed. Although modern file systems (such as NTFS) have a certain degree of self-optimization ability, severe fragmentation will still affect performance.
[0005] Resource consumption: A large number of small-sized files will occupy more memory and CPU resources because each file requires a separate metadata record. In addition, the file system cache may not be effectively utilized, resulting in a higher I / O request rate.
[0006] Difficult backup and recovery: A dataset containing a large number of small-sized files may be very time-consuming during backup and recovery because each file must be processed separately. This not only increases the time cost during the backup process but also makes disaster recovery more complex.
[0007] Space waste: Since the file system usually allocates a certain minimum storage unit (cluster) for each file, if the file size is smaller than this unit, the remaining space will be wasted. For a large number of small-sized files, this space waste may accumulate to a considerable extent.
[0008] Nowadays, the popularity of short videos, WeChat, digital cameras, mobile phone cameras, etc. has led to the generation of a large number of small-sized files in computers. When the existing file management system manages such files, the buffer efficiency during file reading drops rapidly. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a virtual file system suitable for high-speed access to a batch of small-sized files. The small-sized files mentioned in this article generally refer to files with a size not exceeding 50MB.
[0010] The specific technical solution for the present invention to solve the above technical problems is as follows:
[0011] A virtual file system applicable to high-speed access of a batch of small-sized files, including a virtual file body and a virtual file reading and writing program;
[0012] The virtual file body contains a number of small-sized files, and the small-sized files refer to various photo files and video files;
[0013] The virtual file body includes a file header, a file index table, and a file data body;
[0014] The file header includes a file flag indicating that it is a virtual file, the start address and end address of the virtual file body;
[0015] The file data body is a small-sized file body stored in sequence;
[0016] The file index table includes information about the small-sized files stored in the file data body, and the information of the small-sized files includes the small-sized file name, the start storage address of the small-sized file, the end storage address of the small-sized file, the attribute remarks of the small-sized file, the size of the small-sized file, the small-sized file category, and the thumbnail data of the small-sized file;
[0017] The virtual file reading and writing program is used to retrieve, read, and delete the small-sized files contained in the virtual file body, add small-sized files to the virtual file body, and sort and rename the small-sized files contained in the virtual file body.
[0018] Further, the file header further includes the total number of small-sized files contained in the virtual file body and the size of the virtual file body. The size of the virtual file body, when continuously stored, is the difference between the start address and end address of the virtual file body. However, since the virtual file body may also be stored dispersedly, when stored dispersedly, the start address of the virtual file body is the minimum address in the physical storage space, and the end address of the virtual file body is the maximum address in the physical storage space. At this time, the size of the virtual file body is not simply the difference between the start address and end address, but the sum of the sizes of all small-sized files in the virtual file body, the size of the file header of the virtual file body, and the size of the file index table.
[0019] Further, when the virtual file reading and writing program retrieves the small-sized files contained in the virtual file body, it only retrieves the information in the file index table and does not retrieve the file data body;
[0020] Further, when the virtual file reading and writing program sorts small-sized files contained in the virtual file body, it only sorts the information in the file index table and does not modify the data in the file data body. As for the retrieval rules, they can follow the existing technologies. For example, sorting can be in ascending or descending order by name. Of course, specific retrieval factors can be adjusted according to specific situations by adding or subtracting information in the file retrieval table, such as adding creation time, file size, etc. The retrieval rules themselves are not the innovation points of the present invention. The innovation point of the present invention is that the search range of the system in the physical storage space during retrieval is limited to the storage space where the file index table is located, thereby saving search time. Since the file index table includes the storage location of each small file in the file data body, no matter how the sorting is done, as long as the file index table can correctly point to the data location in the file data body, so during sorting, only the information in the file index table is sorted, and there is no need to modify the data in the file data body.
[0021] Further, when the virtual file reading and writing program adds a small-sized file to the virtual file body, it includes the following steps:
[0022] S11 Update the parameter "total number of included files" in the file header;
[0023] S12 Append the information of the newly added small-sized file to the end of the existing data in the index table;
[0024] S13 Append the data of the small-sized file to the file data body.
[0025] Further, when the virtual file reading and writing program deletes a certain small-sized file from the virtual file body, it includes the following steps:
[0026] S21 Update the parameter "total number of included files" in the file header;
[0027] S22 Set the corresponding item of the small-sized file to be deleted in the index table in the virtual file body to be empty;
[0028] S23 Set the storage space corresponding to the file data body of the small-sized file to be deleted in the virtual file body to be free space. However, at this time, the data block space of the deleted small-sized file has not been released from the virtual file to the hard disk free space, so at this time, the size of the virtual file body has not become smaller, but the sum of the sizes of all the included small-sized files has become smaller due to the deletion operation. That is, the size of the virtual file body, more specifically, also includes the size of the small-sized files deleted between two file optimization operations, which leads to the concept of the "usage rate index".
[0029] Further, the file header further includes a usage rate index, where the usage rate index = (sum of the sizes of all small-sized files included) / (size of the virtual file body), and when the usage rate index is less than 90%, file optimization is initiated.
[0030] Further, the file optimization refers to defragmentation work, which stores the virtual file body in continuous storage space, releases the space of the deleted small-sized files to the free space of the memory, and updates various start addresses and end addresses in the virtual file body. Optimize the file header of the virtual file, and delete the information of the small-sized files set to NULL (i.e., deleted) from the index table.
[0031] Further, the small-sized files are multimedia files, such as pictures, photos, short videos, etc.
[0032] The present solution has the following beneficial effects compared with the prior art:
[0033] Through the above virtual file system, the present solution can store a large number of small-sized files in a large virtual file. For subsequent operations such as sorting and retrieving a large number of small-sized files, only the index table is retrieved, avoiding the situation of global retrieval in the storage device, and improving the retrieval efficiency and batch deletion efficiency. After some files in the virtual file are deleted, calculate the proportion of free space in a timely manner. If the proportion of free space in the entire virtual file size exceeds 10%, space optimization will be performed on the virtual file. The present invention can improve the batch processing efficiency of photo files, short video files and other similar large numbers of small-sized files. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a structural diagram of the virtual file body of the present invention;
[0035] Figure 2 It is a read-write flow chart of the virtual file reading program of the present invention;
[0036] Figure 3 It is a comparison table of the existing storage method and the virtual file storage method based on the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0038] A virtual file system suitable for high-speed access to a batch of small-sized files includes a virtual file body and a virtual file read-write program;
[0039] The virtual file body contains several small-sized files, and the small-sized files refer to various photo files and video files. In this example, traditional photo files are used for storage:
[0040] As Figure 1 shown, the virtual file body includes a file header, a file index table, and a file data body;
[0041] The file header includes a file flag, the start address and end address of the virtual file body, the size of the virtual file body, and the file usage rate; it also includes the total number of photo files contained in the virtual file body and a flag bit indicating whether optimization is required.
[0042] The file flag is used to distinguish whether this file is a virtual file. In this example, it is set to "VF", which means a virtual file. As Figure 2 shown, to let the existing file system know that this file needs to be operated by a virtual file reading and writing program. The start address and end address of the virtual file body refer to the minimum address and maximum address occupied by the entire virtual file in the physical storage space. The usage rate index = (the sum of the sizes of all included photos) / (the size of the virtual file body). When the usage rate index is less than 90%, the file optimization flag is started, and defragmentation is performed when the computer system is in an idle time.
[0043] The file data body is the original data of the photo files stored in sequence;
[0044] The index table includes information about the photos stored in the file data body. The information about the photos includes a file sequence, the start address where the original file of the photo is stored, the end address where the original file of the photo is stored, the file name of the photo, the size of the photo file, the file category of the photo, and its thumbnail data;
[0045] The file sequence is which file, such as Figure 1 shown as "Photo File 1", "Photo File 2", etc. in
[0046] The start address where the original file of the photo is stored and the end address where the original file of the photo is stored refer to the start position and end position recorded by the photo file in the file data body of this virtual file. Here, only the actual storage position of the original data of the photo file is recorded, and the original data itself is not recorded, so the occupied space is extremely small;
[0047] The file name of the photo and the size of the photo file do not need much explanation;
[0048] The categories of photos can be in various different formats, such as JPG, PNG, RAW, etc. Multiple format files can be stored in a virtual file. However, generally, it is used to store files of the same type, such as all storing pictures, or all storing video files, etc. Generally, video files and photo files are not mixed together. Although theoretically it is also possible to store them in this way, for the actual applications of users, such storage has no advantage in improving the read and write efficiency.
[0049] The thumbnail information corresponds to the thumbnail of the photo file. The information contained in the thumbnail is much smaller than the original data. When performing retrieval and sorting, only the thumbnail information is required, so there is no need to read the original data of the picture. Since the file retrieval table contains the start address and end address where the small file is stored in the "file data body", when re - sorting the small file, only the order in the file index table needs to be adjusted. The storage address information corresponding to the small file will be moved together, and there is no need to move the data in the small file data body. Therefore, the amount of data that needs to be read and written during sorting is also greatly reduced.
[0050] The virtual file read - write program is used to retrieve, read, and delete the photo files contained in the virtual file body, add photo files to the virtual file body, re - sort, rename, etc. the photo files contained in the virtual file body.
[0051] That is, on the existing file system, a virtual file read - write program is built, and a virtual file format is built to integrate a large number of photo files and store them in a virtual file.
[0052] When the virtual file read - write program retrieves and sorts the photo files contained in the virtual file body, it only operates in the file index table and does not retrieve and sort the information of the file data body.
[0053] Such as Figure 3 shown, it is a comparison diagram of the storage method of photo files under the existing file system and the storage method of photo files in this virtual file system.
[0054] When the virtual file read - write program adds a photo file to the virtual file body, it includes the following steps:
[0055] S11 Update the parameter "total number of included files" in the file header;
[0056] S12 Append the information of the newly added small - size file to the end of the existing data in the index table;
[0057] S13 Append the data of the small - size file to the file data body.
[0058] A specific case is as follows:
[0059] If the virtual file does not exist, create a new virtual file
[0060] Set the file flag bit = VF
[0061] Obtain the starting address of the free storage space, reserve the storage space for the file header and the file index table, and set the starting address of the file data body;
[0062] Set the number of photo files in the virtual file header = the number of photo files + 1;
[0063] Set the photo file.startAddress in the index table = StartAddress
[0064] Set the photo file.Lenth in the index table = FileLenth
[0065] Set the photo file.EndAddress in the index table = StartAddress + FileLenth
[0066] Set the photo file(1).type in the index table = JPG;
[0067] Set the starting address of the virtual file to: the starting address of the free storage space;
[0068] Set the ending address of the virtual file to: the original virtual file ending address + the total length of the newly written photo files; if the virtual file is newly created, the original virtual file ending address = the starting address of the free storage space + the reserved file header space + the reserved file index table space;
[0069] Calculate the virtual file usage rate = the sum of the photo files contained in the virtual file / the size of this virtual file;
[0070] If the virtual file usage rate is less than 90%, set the optimization flag bit in the file header to 1, otherwise set it to 0;
[0071] End.
[0072] If the optimization flag of the virtual file is 1, the system starts to defragment the virtual file during idle time.
[0073] Since the virtual file itself needs to occupy a file header and an index table of several hundred KB to several MB (depending on the total number and size of small-sized files), generally when the available physical storage space is not less than 2G, creating and using this virtual file will show the superiority of the efficiency of this solution.
[0074] The reason why the relevant data is appended at the end of the existing data in the index in step S12 is that the space of the index table is a pre-allocated storage space when the virtual file body is established; while in step S13, it only mentions appending relevant data in the file data body because the appended data may be at the end of the existing data, or it may be in an idle space vacated due to file deletion in the middle.
[0075] Further, when the virtual file reading and writing program deletes a certain photo file from the virtual file body, it includes the following steps:
[0076] S21 Update the "total number of included files" in the file header, that is, the new value of the total number of included files = the old value of the total number of included files - 1;
[0077] S22 Set the corresponding item of the small-sized file to be deleted in the index table in the virtual file body to be empty;
[0078] S23 Set the storage space corresponding to the small-sized file to be deleted in the file data body of the virtual file body to be idle space.
[0079] Specific examples are as follows:
[0080] S21. Obtain the serial number of the photo file in the virtual file: GetFileIndex(num);
[0081] S21-1. Calculate the usage rate of the virtual file:
[0082] CntUsingRate = (SumLenth - FileIndex.Num(num).Lenth) / SumLenth
[0083] S22. Set the corresponding position in the virtual file index table to be empty: SetFileIndex.Num(num) = NULL;
[0084] S23. Release the corresponding storage space in the file data body:
[0085] Release(FileIndex.Num(num); GetFileIndex(num);
[0086] As Figure 1 shown, photo file 3 has been deleted, so the corresponding data in the file index table is all empty, and the corresponding storage space in the file data body is idle storage space.
[0087] It should be noted that, under ideal conditions, the storage of the virtual file body in the storage device is continuous, that is, in the file data body, the end address of photo file 1 is adjacent to the start address of photo file 2; when a new photo file is inserted, the virtual file body will preferentially store behind the last photo file, and only when the physical storage space is exhausted will it look for the storage space vacated by the deleted photo file in front. This strategy can minimize the search time for free space. On the other hand, the virtual file will recalculate the usage rate of the virtual file (which can be completed only through the index table information) after each addition or deletion of a photo file. Once the usage rate is lower than 90%, it will perform defragmentation during the idle time of the system.
[0088] Furthermore, the file optimization refers to the defragmentation work to make the virtual file body stored in continuous storage space.
[0089] Based on the above virtual file system, the test read time is as follows:
[0090] File test object: Photo files, with sizes ranging from 3MB to 10MB, totaling 512MB, and a total of 100 photo files, which are written into a virtual file. The test environment is the Linux system.
[0091] The specific test data is shown in Table 1:
[0092]
[0093] Table 1
[0094] The statistical data is shown in Table 2 below:
[0095]
[0096] Table 2
[0097] Because the writing process is slower compared to the read operation, only the comparison results of the writing time are listed here: When the traditional file management system performs a write operation, the worst value is 74.7 ms / MB, while in the above experiment, the worst value of the write operation using this virtual file system is 51 ms / MB.
[0098] From the above test data, it can be seen that using this virtual file system can improve the access speed of a large number of small-size data and optimize the user experience.
[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A virtual file system suitable for high-speed access to batches of small-size files, which is characterized in that it includes a virtual file body and a virtual file reading and writing program; The virtual file body contains a number of small-size files; The virtual file body includes a file header, a file index table and a file data body; The file header includes a file mark, a start address and an end address of the virtual file body; The file data body is a small-size file body stored sequentially; The file index table includes information of small-sized files stored in the file data body, wherein the information of small-sized files includes small-sized file names, storage start addresses of small-sized files, storage end addresses of small-sized files, small-sized file categories, and thumbnail data of small-sized files; The virtual file reading and writing program is used to retrieve, read, delete small-sized files contained in the virtual file body, add small-sized files to the virtual file body, and sort and rename small-sized files contained in the virtual file body.
2. The virtual file system suitable for high-speed access to batches of small-size files according to claim 1, characterized in that: The file header also includes the total number of small-size files contained in the virtual file body and the size of the virtual file body.
3. The virtual file system suitable for high-speed access to batches of small-size files according to claim 2, characterized in that: When retrieving a small-size file contained in the virtual file body, the virtual file reading and writing program only retrieves the file index table information but does not retrieve the file data body.
4. The virtual file system suitable for high-speed access to batches of small-size files according to any one of claims 1 to 3, characterized in that: When the virtual file reading and writing program sorts the small-size files contained in the virtual file body, it only sorts the information in the file index table and does not change the data in the file data body.
5. The virtual file system suitable for high-speed access to batches of small-size files according to any one of claims 1 to 3, characterized in that: When the virtual file reading and writing program adds a small-size file to the virtual file body, the following steps are included: S11, update the "total number of included files" parameter in the file header; S12, adding information of newly added small-size files to the end of the existing data in the index table; S13. Add the data of the small-size file to the file data body.
6. The virtual file system suitable for high-speed access to batches of small-size files according to claim 5, characterized in that: When the virtual file reading and writing program deletes a small-size file from the virtual file body, the following steps are included: S21, update the "total number of included files" parameter in the file header; S22, setting the corresponding item of the small-size file to be deleted in the index table in the virtual file body to be empty; S23: Set the storage space corresponding to the file data body of the to-be-deleted small-size file in the virtual file body as free space.
7. The virtual file system suitable for high-speed access to batches of small-size files according to any one of claims 1 to 3, characterized in that: The file header also includes a usage rate indicator, where the usage rate indicator = (the sum of the sizes of all small-sized files included) / (the size of the virtual file body), and when the usage rate indicator is less than 90%, file optimization is started.
8. The virtual file system suitable for high-speed access to batches of small-size files according to any one of claims 1 to 3, characterized in that: The file optimization refers to defragmentation work so that the virtual file body is stored in a continuous storage space.
9. The virtual file system suitable for high-speed access to batches of small-size files according to any one of claims 1 to 3, characterized in that: The small-size file is a multimedia file.