Intelligent small file throughput optimization method and system

Through aggregation of small files, distributed metadata storage, dynamic storage adjustment and real-time optimization, the problems of high metadata pressure, low storage efficiency and high access latency in small files processing are solved, and the performance and user experience of the storage system are significantly improved.

CN120010780AInactive Publication Date: 2025-05-16BEIJING RONGXUN OPTICAL TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510098972.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing small files, the prior art has high pressure on metadata management, low storage efficiency and high access latency, which seriously affects the performance and user experience of the storage system.

Method used

By aggregating multiple small files into logical files, using distributed metadata storage and indexing mechanisms, the access frequency and importance of small files are obtained, the storage location is dynamically adjusted, and real-time optimization and adjustment are carried out.

Benefits of technology

It significantly reduces the complexity and pressure of metadata management, improves the storage efficiency and access performance of small files, reduces the amount of metadata, optimizes the utilization of storage space, and reduces the access delay of small files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010780A_ABST
    Figure CN120010780A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of small file processing, and provides an intelligent small file throughput optimization method and system. The method comprises the following steps: aggregating a plurality of small files into a logic file; a distributed metadata storage and indexing mechanism is adopted to store and index a plurality of small files; obtaining the access frequency and importance of the plurality of small files, and storing the plurality of small files in different storage media; and carrying out real-time load and access monitoring, dynamically adjusting the storage positions of the plurality of small files, and carrying out optimization adjustment on aggregation of the plurality of small files. The complexity and pressure of metadata management are remarkably reduced, and the storage efficiency and access performance of small files are fully improved, so that the bottleneck of the prior art is broken, the metadata volume can be effectively reduced, the utilization of the storage space is ingeniously optimized, and the access delay of the small files is greatly reduced. And the small file processing capability of the whole cluster storage system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of small file processing, and in particular relates to an intelligent small file throughput optimization method and system. Background Art

[0002] In today's era of rapid digitalization, the demand for data storage has grown rapidly, and cluster storage products have been widely and deeply applied in many key fields such as data centers and cloud computing. With the vigorous development of various businesses, efficient processing of small files has gradually become a major challenge facing storage systems. Small files usually refer to files with a size ranging from a few KB to hundreds of KB, such as web files, image files, log files, etc. In a large-scale storage environment, the number of small files is vast, and the efficiency of their management and access has a significant impact on the performance of the entire storage system.

[0003] The distributed storage architecture improves the reliability and scalability of the system by storing data in multiple nodes; the data redundancy strategy ensures the security and availability of data, and data will not be lost even if some nodes fail; the data index mechanism is used to quickly locate and retrieve data, improving the efficiency of data access.

[0004] In the prior art, cluster storage products generally adopt traditional file systems and storage architectures when dealing with small files. For example, based on traditional distributed file systems (such as HDFS), small files are directly stored in the file system, and the files are searched and accessed with the help of centralized metadata management. Therefore, the prior art has the following defects: (1) High metadata management pressure: The metadata of a large number of small files are centrally managed, which makes the metadata server overwhelmed and overloaded, which seriously restricts the scalability and performance of the system; (2) Low storage efficiency: The separate storage of small files inevitably causes a huge waste of storage space, and the frequent disk I / O operations are like adding insult to injury, greatly reducing storage efficiency; (3) High access delay: Due to the decentralized storage characteristics of small files and the huge query overhead of metadata, the access delay of small files remains high, which seriously affects the user experience. Summary of the invention

[0005] The purpose of the embodiments of the present invention is to provide an intelligent small file throughput optimization method and system, aiming to solve the technical problems existing in the prior art mentioned in the background technology.

[0006] The embodiment of the present invention is implemented as follows:

[0007] An intelligent small file throughput optimization method, the method specifically comprising the following steps:

[0008] Aggregate multiple small files into logical files;

[0009] Use a distributed metadata storage and indexing mechanism to store and index multiple small files;

[0010] Obtaining access frequencies and importances of the plurality of small files, and storing the plurality of small files on different storage media;

[0011] Real-time load and access monitoring is performed, the storage locations of the plurality of small files are dynamically adjusted, and the aggregation of the plurality of small files is optimized and adjusted.

[0012] As a further limitation of the technical solution of the embodiment of the present invention, the aggregating multiple small files into a logical file specifically includes the following steps:

[0013] Receive user's file aggregation settings;

[0014] According to the file aggregation setting, determine the number of small file aggregations;

[0015] The small files are aggregated into logical files according to the small file aggregation quantity.

[0016] As a further limitation of the technical solution of the embodiment of the present invention, the aggregating multiple small files into a logical file specifically includes the following steps:

[0017] Receive user's file aggregation settings;

[0018] According to the file aggregation setting, determining a plurality of similar aggregation characteristics;

[0019] According to the multiple similar aggregation characteristics, multiple small files are classified and aggregated to obtain multiple logical files.

[0020] As a further limitation of the technical solution of the embodiment of the present invention, the use of a distributed metadata storage and indexing mechanism to store and index multiple small files specifically includes the following steps:

[0021] A distributed metadata storage and indexing mechanism is used to perform storage and index planning for the plurality of small files, and the planning results are recorded;

[0022] According to the planning result, multiple small files are stored and indexed.

[0023] As a further limitation of the technical solution of the embodiment of the present invention, the use of a distributed metadata storage and indexing mechanism to store and index multiple small files specifically includes the following steps:

[0024] Using a distributed metadata storage and indexing mechanism, a blockchain storage plan is performed for the plurality of small files, and the planning results are recorded;

[0025] According to the planning result, multiple small files are stored and indexed.

[0026] As a further limitation of the technical solution of the embodiment of the present invention, obtaining the access frequency and importance of the plurality of small files and storing the plurality of small files on different storage media specifically includes the following steps:

[0027] Obtaining access frequency and importance of multiple small files;

[0028] Based on a preset frequency threshold and importance weight, storage analysis is performed on the access frequency and importance of the plurality of small files, and the corresponding storage media are matched;

[0029] The plurality of small files are stored in a corresponding storage medium.

[0030] As a further limitation of the technical solution of the embodiment of the present invention, the real-time load and access monitoring, dynamic adjustment of the storage locations of the plurality of the small files, and optimization adjustment of the aggregation of the plurality of the small files specifically include the following steps:

[0031] Conduct real-time load and access monitoring to determine whether dynamic adjustments are needed;

[0032] When dynamic adjustment is required, the storage locations of the plurality of small files are dynamically adjusted, and the aggregation of the plurality of small files is optimized and adjusted.

[0033] An intelligent small file throughput optimization system, the system includes a small file aggregation module, a metadata management module, a hierarchical storage module and an intelligent scheduling module, wherein:

[0034] Small file aggregation module, used to aggregate multiple small files into a logical file;

[0035] The metadata management module is used to store and index multiple small files using a distributed metadata storage and indexing mechanism;

[0036] A hierarchical storage module, used for obtaining the access frequency and importance of the plurality of small files, and storing the plurality of small files on different storage media;

[0037] The intelligent scheduling module is used to perform real-time load and access monitoring, dynamically adjust the storage locations of multiple small files, and optimize the aggregation of multiple small files.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The embodiment of the present invention aggregates multiple small files into a logical file; adopts a distributed metadata storage and indexing mechanism to store and index multiple small files; obtains the access frequency and importance of multiple small files, and stores multiple small files on different storage media; performs real-time load and access monitoring, dynamically adjusts the storage location of multiple small files, and optimizes the aggregation of multiple small files. It significantly reduces the complexity and pressure of metadata management, fully improves the storage efficiency and access performance of small files, thereby breaking the bottleneck of existing technologies, effectively reducing the amount of metadata, subtly optimizing the use of storage space, significantly reducing the access delay of small files, and improving the processing capacity of the entire cluster storage system for small files. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flow chart of an intelligent small file throughput optimization method provided by an embodiment of the present invention is shown;

[0041] Figure 2 The application architecture diagram of the intelligent small file throughput optimization system provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0043] It is understandable that in the prior art, cluster storage products generally adopt traditional file systems and storage architectures when dealing with small files. For example, based on traditional distributed file systems (such as HDFS), small files are directly stored in the file system, and the files are searched and accessed with the help of centralized metadata management. Therefore, the prior art has the following defects: (1) High metadata management pressure: The metadata of a large number of small files are centrally managed, which makes the metadata server overwhelmed and overloaded, which seriously restricts the scalability and performance of the system; (2) Low storage efficiency: The way of storing small files separately inevitably causes a huge waste of storage space, and the frequent disk I / O operations are like adding insult to injury, greatly reducing storage efficiency; (3) High access delay: Due to the decentralized storage characteristics of small files and the huge query overhead of metadata, the access delay of small files remains high, which seriously affects the user experience.

[0044] To solve the above problems, an intelligent small file throughput optimization method and system disclosed in an embodiment of the present invention aggregates multiple small files into logical files; uses a distributed metadata storage and indexing mechanism to store and index multiple small files; obtains the access frequency and importance of multiple small files, and stores multiple small files on different storage media; performs real-time load and access monitoring, dynamically adjusts the storage location of multiple small files, and optimizes the aggregation of multiple small files. It significantly reduces the complexity and pressure of metadata management, fully improves the storage efficiency and access performance of small files, thereby breaking the bottleneck of the existing technology, effectively reducing the amount of metadata, subtly optimizing the use of storage space, greatly reducing the access delay of small files, and improving the processing capacity of the entire cluster storage system for small files.

[0045] Specifically, Figure 1 A flow chart of an intelligent small file throughput optimization method provided by an embodiment of the present invention is shown.

[0046] In a preferred embodiment of the present invention, an intelligent small file throughput optimization method is provided, and the method specifically comprises the following steps:

[0047] Step S101: Aggregate multiple small files into a logical file.

[0048] In an embodiment of the present invention, a user can make different file aggregation settings. By receiving the user's file aggregation settings, specifically: in one case, according to the file aggregation settings, the number of small file aggregations is determined, and then the small files are aggregated into logical files according to the number of small file aggregations; in another case, according to the file aggregation settings, multiple similar aggregation characteristics are determined, and then according to the multiple similar aggregation characteristics, multiple small files are classified and aggregated to obtain multiple logical files, thereby effectively reducing the number of files and greatly reducing the amount of metadata.

[0049] Step S102: Use a distributed metadata storage and indexing mechanism to store and index multiple small files.

[0050] In an embodiment of the present invention, by adopting a distributed metadata storage and indexing mechanism, storage and indexing planning is performed on multiple small files, and the planning results are recorded, or blockchain storage planning is performed on multiple small files, and the planning results are recorded, and then according to the planning results, multiple small files are stored and indexed, thereby improving the query efficiency and scalability of metadata.

[0051] Step S103: Obtain access frequencies and importances of the multiple small files, and store the multiple small files on different storage media.

[0052] In an embodiment of the present invention, by obtaining the access frequency and importance of multiple small files, storage analysis is performed on the access frequency and importance of multiple small files based on preset frequency thresholds and importance weights, and corresponding storage media are matched. Then, multiple small files are stored in corresponding storage media, and the migration conditions of files between different storage media are determined to achieve reasonable allocation of resources, such as high-speed SSDs and large-capacity HDDs, to achieve optimal configuration of storage resources.

[0053] Step S104: perform real-time load and access monitoring, dynamically adjust the storage locations of the plurality of small files, and optimize the aggregation of the plurality of small files.

[0054] In an embodiment of the present invention, real-time load and access monitoring is performed to determine whether dynamic adjustment is required. If it is determined that dynamic adjustment is required, the storage locations of multiple small files are dynamically adjusted, and the aggregation of multiple small files is optimized to ensure that the system is always in an efficient operating state.

[0055] Furthermore, Figure 2 The application architecture diagram of the intelligent small file throughput optimization system provided by an embodiment of the present invention is shown.

[0056] Among them, in another preferred embodiment provided by the present invention, an intelligent small file throughput optimization system includes:

[0057] The small file aggregation module 101 is used to aggregate multiple small files into a logical file.

[0058] In an embodiment of the present invention, a user may make different file aggregation settings. The small file aggregation module 101 receives the file aggregation settings of the user. Specifically: in one case, the number of small file aggregations is determined according to the file aggregation settings, and then the small files are aggregated into logical files according to the number of small file aggregations; in another case, a plurality of similar aggregation characteristics are determined according to the file aggregation settings, and then the plurality of small files are classified and aggregated according to the plurality of similar aggregation characteristics to obtain a plurality of logical files, thereby effectively reducing the number of files and greatly reducing the amount of metadata.

[0059] The metadata management module 102 is used to store and index multiple small files using a distributed metadata storage and indexing mechanism.

[0060] In an embodiment of the present invention, the metadata management module 102 adopts a distributed metadata storage and indexing mechanism to perform storage and indexing planning for multiple small files, record the planning results, or perform blockchain storage planning for multiple small files, record the planning results, and then store and index multiple small files according to the planning results, thereby improving the query efficiency and scalability of metadata.

[0061] The hierarchical storage module 103 is used to obtain the access frequency and importance of the multiple small files and store the multiple small files on different storage media.

[0062] In the embodiment of the present invention, the hierarchical storage module 103 obtains the access frequency and importance of multiple small files, performs storage analysis on the access frequency and importance of multiple small files based on preset frequency thresholds and importance weights, matches corresponding storage media, and then stores multiple small files in corresponding storage media, determines the migration conditions of files between different storage media, and realizes reasonable allocation of resources, such as high-speed SSD and large-capacity HDD, to achieve optimal configuration of storage resources.

[0063] The intelligent scheduling module 104 is used to perform real-time load and access monitoring, dynamically adjust the storage locations of the plurality of small files, and optimize the aggregation of the plurality of small files.

[0064] In an embodiment of the present invention, the intelligent scheduling module 104 performs real-time load and access monitoring to determine whether dynamic adjustment is required. If it is determined that dynamic adjustment is required, the storage locations of multiple small files are dynamically adjusted, and the aggregation of multiple small files is optimized to ensure that the system is always in an efficient operating state.

[0065] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0066] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0067] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. An intelligent small file throughput optimization method, characterized in that: The method specifically comprises the following steps: Aggregate multiple small files into logical files; Use a distributed metadata storage and indexing mechanism to store and index multiple small files; Obtaining access frequencies and importances of the plurality of small files, and storing the plurality of small files on different storage media; Real-time load and access monitoring is performed, the storage locations of the plurality of small files are dynamically adjusted, and the aggregation of the plurality of small files is optimized and adjusted.

2. The intelligent small file throughput optimization method according to claim 1 is characterized in that: Aggregating multiple small files into a logical file specifically includes the following steps: Receive user's file aggregation settings; According to the file aggregation settings, determine the number of small file aggregations; The small files are aggregated into logical files according to the small file aggregation quantity.

3. The intelligent small file throughput optimization method according to claim 1 is characterized in that: Aggregating multiple small files into a logical file specifically includes the following steps: Receive user's file aggregation settings; According to the file aggregation setting, determining a plurality of similar aggregation characteristics; According to the multiple similar aggregation characteristics, multiple small files are classified and aggregated to obtain multiple logical files.

4. The intelligent small file throughput optimization method according to claim 1, characterized in that: The distributed metadata storage and indexing mechanism is used to store and index multiple small files, and specifically includes the following steps: A distributed metadata storage and indexing mechanism is used to perform storage and index planning for the plurality of small files, and the planning results are recorded; According to the planning result, multiple small files are stored and indexed.

5. The intelligent small file throughput optimization method according to claim 1 is characterized in that: The distributed metadata storage and indexing mechanism is used to store and index multiple small files, and specifically includes the following steps: Using a distributed metadata storage and indexing mechanism, a blockchain storage plan is performed for the plurality of small files, and the planning results are recorded; According to the planning result, multiple small files are stored and indexed.

6. The intelligent small file throughput optimization method according to claim 1 is characterized in that: The obtaining of the access frequency and importance of the plurality of small files and storing the plurality of small files on different storage media specifically comprises the following steps: Obtaining access frequency and importance of multiple small files; Based on a preset frequency threshold and importance weight, storage analysis is performed on the access frequency and importance of the plurality of small files, and the corresponding storage media are matched; The plurality of small files are stored in a corresponding storage medium.

7. The intelligent small file throughput optimization method according to claim 1, characterized in that: The real-time load and access monitoring, dynamic adjustment of the storage locations of the plurality of small files, and optimization adjustment of the aggregation of the plurality of small files specifically include the following steps: Conduct real-time load and access monitoring to determine whether dynamic adjustments are needed; When dynamic adjustment is required, the storage locations of the plurality of small files are dynamically adjusted, and the aggregation of the plurality of small files is optimized and adjusted.

8. An intelligent small file throughput optimization system, characterized in that: The system includes a small file aggregation module, a metadata management module, a hierarchical storage module and an intelligent scheduling module, wherein: Small file aggregation module, used to aggregate multiple small files into a logical file; The metadata management module is used to store and index multiple small files using a distributed metadata storage and indexing mechanism; A hierarchical storage module, used for obtaining the access frequency and importance of the plurality of small files, and storing the plurality of small files on different storage media; The intelligent scheduling module is used to perform real-time load and access monitoring, dynamically adjust the storage locations of multiple small files, and optimize the aggregation of multiple small files.

Citation Information

Cited By

  • Data optimization storage method and device based on artificial intelligence and medium

    CN121349968A