Distributed high-concurrency aggregation storage method and system for small files
By calculating the real-time CPU processing speed and file reading concurrency, selecting the target server, compressing the text file, selecting the storage directory using file name hash value and polling count value, and aggregating small files to store, solving the efficiency and performance problems of the small file storage system under high concurrent access, and achieving efficient file retrieval and access.
Patent Information
- Application Number
- CN202510456812.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-12
- Publication Date
- 2025-07-11
AI Technical Summary
In the high concurrent access scenario, it is difficult for the existing technology to effectively optimize the distributed storage of small files, resulting in low storage space utilization, degraded file system performance and inefficient retrieval efficiency, and lack of real-time monitoring and dynamic scheduling mechanisms for resource utilization of storage nodes.
By obtaining the file type and size, the real-time CPU processing speed is calculated based on the CPU main frequency and historical data processing volume, the target server is selected based on the queue rate, the concurrency of file reading and the remaining space of the disk, and the text file is compressed, the storage directory is selected using the file name hash value and polling count value, and multiple small files are aggregated in the same target aggregate file, a file name containing storage location information is generated, and the data block is located through the index table when receiving the file access request.
Improve storage efficiency, reduce file system load, improve file retrieval efficiency and access performance, avoid waste of storage space, and ensure the fastness and accuracy of file access.
Smart Images

Figure CN120295977A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of file storage, and in particular to a distributed high-concurrency aggregate storage method and system for small files. Background Art
[0002] With the rapid development of Internet technology and the continuous expansion of data scale, small file storage has become a technical problem that needs to be solved urgently. In practical applications, the storage of a large number of small files will not only lead to inefficient use of storage space, but also cause file system performance degradation and low file retrieval efficiency. Especially in high-concurrency access scenarios, frequent small file read and write operations will bring huge pressure to the storage system, and also increase the complexity of system maintenance and management.
[0003] At present, the main solution adopted by the industry is to store small files in a package, that is, to merge multiple small files into a large file for storage, and record the location information of each small file in the large file through an index table. This method can effectively reduce the index node occupation of the file system, improve the utilization of storage space, and also improve file access performance to a certain extent.
[0004] However, when the system needs to perform load balancing, due to the lack of real-time monitoring and dynamic scheduling mechanism of storage node resource utilization, it is difficult to optimize storage allocation according to the load conditions of each node. This makes it difficult to fully utilize the overall performance of the system in a distributed environment, especially in high-concurrency access scenarios. Summary of the invention
[0005] The present application provides a small file distributed high-concurrency aggregate storage method and system for optimizing the efficiency and effect of small file storage.
[0006] In a first aspect, the present application provides a small file distributed high-concurrency aggregate storage method, which obtains the file type and file size of the file to be stored uploaded by the client; Each file server collects performance indicator data based on a preset time interval and sends it to the resource coordination server. The performance indicator data includes the real-time CPU processing speed, file processing queue data size, disk remaining space, and file reading concurrency calculated based on the initial benchmark value of the CPU main frequency combined with the historical data processing volume; The queuing rate of each file server calculated by the receiving resource coordination server is the ratio of the file processing queue data size to the real-time CPU processing speed; According to the queuing rate, file reading concurrency and disk remaining space, the files are sorted in descending order of queuing rate, descending order of file reading concurrency and ascending order of disk remaining space, and the file server with the best sorting result is selected as the target server; When the file type of the file to be stored is a text type, the file to be stored is compressed to obtain a compressed file and a compression flag; when the file type of the file to be stored is a non-text type, the file to be stored and an uncompressed flag are used as the file to be written. Obtain the current polling count value in the target server, and perform a hash operation on the file name of the file to be stored to obtain a first-level directory hash value. Select a first-level storage directory according to the first-level directory hash value, select a second-level storage directory according to the polling count value, and increment the polling count value cyclically. Create a target aggregated file in the second-level storage directory, generate a 2-byte file format field according to the compression flag or uncompressed flag, and generate a 4-byte file length field according to the size of the file to be written. Write the file format field, the file length field, and the file to be written into the target aggregated file in sequence, and record the starting offset of the write. Generate a file name containing storage location information according to the server identifier of the target server, the first-level directory hash value, the polling count value, the target aggregated file identifier, and the starting offset, and return the file name to the client.
[0007] By adopting the above technical solutions, by calculating the real-time CPU processing speed based on the initial reference value of the CPU main frequency and the historical data processing volume, combining the size of the file processing queuing data to obtain the queuing rate, and selecting the target server in the priority order of the queuing rate, the file reading concurrency, and the remaining disk space, the actual processing capacity and load status of the server can be accurately reflected, and the file can be prevented from being allocated to a server with poor performance or heavy load. Compressing the text type file can reduce the storage space occupancy and improve the storage efficiency. By using the file name hash value to select the first-level storage directory and the polling count value to select the second-level storage directory, the file distribution is more uniform, the number of files in a single directory is reduced, and the file retrieval efficiency is improved. Aggregating and storing multiple small files in the same target aggregated file reduces the file system load. By recording the compression flag in the file format field, it can be correctly identified whether the file needs to be decompressed when reading. Generating a file name containing complete storage location information enables the specific storage location to be quickly located when accessing the file subsequently, reducing the retrieval overhead.
[0008] Combined with some embodiments of the first aspect, in some embodiments, the real-time CPU processing speed calculated according to the initial reference value of the CPU main frequency combined with the historical data processing volume specifically includes: Obtain the total historical data processing volume and the processing completion time within a preset time window. Divide the total historical data processing volume by the processing completion time to obtain the actual processing speed. The initial reference value of the CPU main frequency and the actual processing speed are weighted and calculated according to a preset weighting coefficient to obtain the real-time CPU processing speed.
[0009] By adopting the above technical solution, the actual processing speed is calculated by obtaining the total amount of historical data processed and the processing completion time within a preset time window, and then weighted calculation is combined with the initial reference value of the CPU main frequency to obtain the real-time CPU processing speed, which can dynamically reflect the data processing performance of the server in the actual operating environment. Since the historical data processing situation is considered, this processing speed can comprehensively reflect multiple influencing factors such as the hardware performance, system load, and network condition of the server, and is more accurate than simply using the CPU main frequency as a performance indicator. By using the weighted calculation method, the influence of historical data and the reference value can be balanced, so that the calculated processing speed can not only reflect the basic processing ability of the server, but also adapt to performance fluctuations, thus providing a more reliable basis for server selection and improving the accuracy of load balancing.
[0010] Combined with some embodiments of the first aspect, in some embodiments, the file to be stored is compressed to obtain a compressed file and a compression identifier, which specifically includes: Obtain the file size of the file to be stored. When the file size is greater than the first preset threshold, divide the file to be stored into multiple data blocks; Perform data feature analysis on each data block to obtain the repeatability, entropy value, and data distribution characteristics of each data block; Select a corresponding candidate compression algorithm set from a preset compression algorithm library according to the data characteristics; Use the compression algorithms in the candidate compression algorithm set to perform compression tests on the data blocks respectively to obtain the compression ratio and compression time of each compression algorithm; Based on a preset weight coefficient, calculate the comprehensive score of each compression algorithm according to the compression ratio and compression time, and select the compression algorithm with the highest score as the target compression algorithm; Use the target compression algorithm to perform compression processing on the data blocks to obtain the compressed data blocks; Use the algorithm identifier of the target compression algorithm as the compression identifier, and combine the compressed data blocks to obtain a compressed file.
[0011] By adopting the above technical solutions, by partitioning large files into data blocks and analyzing the duplication degree, entropy value, and data distribution characteristics of each data block, the characteristics of different data blocks can be identified. Based on the data characteristics, candidate compression algorithms are selected from a preset compression algorithm library, avoiding the performance overhead caused by blindly trying all compression algorithms. By obtaining the compression ratio and compression time through compression tests and calculating the comprehensive score according to the preset weight coefficients to select the optimal compression algorithm, a good balance is achieved between the compression effect and processing efficiency. Using different compression algorithms for data blocks with different characteristics can obtain better compression effects. Recording the compression algorithm identifier as the compression identifier ensures that the correct algorithm can be used during decompression. This adaptive compression processing solution can obtain a high compression ratio while ensuring processing efficiency, effectively reducing the storage space occupation.
[0012] In combination with some embodiments of the first aspect, in some embodiments, after recording the starting offset of the write, the method further includes: Detecting whether the current file size of the target aggregated file reaches a preset capacity threshold; When the current file size reaches the preset capacity threshold, defragmenting the data in the target aggregated file; Deleting the data blocks marked as invalid in the target aggregated file, and storing the valid data blocks continuously again; Updating the offset information in the file name corresponding to the valid data blocks.
[0013] By adopting the above technical solutions, by detecting the size of the target aggregated file and defragmenting it when the preset capacity threshold is reached, invalid data blocks can be cleaned up in time, avoiding waste of storage space. Storing the valid data blocks continuously again reduces file fragmentation and improves the utilization efficiency of storage space. Updating the offset information in the file name ensures that the file content can still be accessed correctly after the data block positions change. This dynamic maintenance mechanism enables the aggregated file to always maintain a good storage state, avoiding the problem of storage space fragmentation caused by the increase in file deletion and update operations. By regular defragmentation, a high storage space utilization rate can be maintained, and at the same time, due to the continuous storage of data blocks, the file reading performance can also be improved.
[0014] In combination with some embodiments of the first aspect, in some embodiments, defragmenting the data in the target aggregated file specifically includes: Constructing a file mapping table to record the starting position, length, and valid status of each data block in the target aggregated file; Traversing the file mapping table in the storage order of the data blocks; Sequentially migrating the data blocks with valid status to the starting position of the file to obtain continuously stored valid data; Updating the position information of each valid data block in the file mapping table.
[0015] By adopting the above technical solution, by constructing a file mapping table to record the detailed information of each data block in the target aggregated file and traversing the file mapping table in the storage order for data block migration, the systematic arrangement of file fragments is realized. During the migration process, valid data blocks are sequentially moved to the starting position of the file, eliminating the fragment intervals in the file and forming a continuous storage space. The continuous storage data distribution reduces the number of disk seek operations and seek time, improving the file reading speed. By updating the position information of the data blocks in the file mapping table, the system can accurately track the latest storage position of each data block, ensuring the accuracy of data access. This method of file fragment arrangement based on the mapping table not only ensures the integrity of data but also improves the utilization efficiency of the storage space, enabling the system to maintain stable storage performance during long-term operation.
[0016] Combined with some embodiments of the first aspect, in some embodiments, when updating the position information of each valid data block in the file mapping table, the method further includes: Establish a file name index table to store the mapping relationship between the file name and the corresponding data block; When a file access request is detected, obtain the target data block information from the file name index table according to the file access request; Read data from the target aggregated file based on the target data block information; Perform decompression processing on the read data.
[0017] By adopting the above technical solution, when a file access request is received, the system can directly locate the target data block information through the index table, and then quickly read the required data from the aggregated file. This indexing mechanism avoids the process of searching one by one among a large number of files, significantly reducing the time complexity of file retrieval. At the same time, the system performs decompression processing on the read data, realizing the unity of compressed storage and fast access. By establishing the file name index table, the system can still ensure the fast retrieval and access of data while maintaining a high compression ratio, effectively balancing the relationship between storage efficiency and access efficiency, and enhancing the overall data processing ability of the system.
[0018] Combined with some embodiments of the first aspect, in some embodiments, establishing a file name index table specifically includes: Resolve the file name into a server identifier, directory information, file identifier, and offset; Locate the target storage location based on the server identifier and directory information; Establish a mapping relationship with the target data block according to the file identifier and offset to obtain the file name index table.
[0019] By adopting the above technical solution, a structured file location mechanism is established by parsing the file name into components such as server identification, directory information, file identification, and offset. The system can accurately locate the target storage location based on the server identification and directory information, and establish a mapping relationship with the target data block through the file identification and offset. This multi-level file name parsing scheme enables the system to accurately locate the physical storage location of each file in a distributed environment, reducing resource consumption during the file retrieval process. The parsed structured information is used to establish a file name index table, providing an efficient file access path for the system and significantly improving the file retrieval efficiency and access performance in the distributed storage system.
[0020] In a second aspect, an embodiment of the present application provides a small file distributed high-concurrency aggregated storage system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, when the above instructions run on the system, causing the above system to execute the method described in the first aspect and any possible implementation manner in the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer program product, when the computer program product runs on the system, causing the system to execute the method described in any possible implementation manner in the first aspect.
[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The present application provides a distributed high-concurrency aggregation storage method for small files. By calculating the real-time CPU processing speed based on the initial benchmark value of the CPU main frequency and the historical data processing volume, the queuing rate is obtained in combination with the file processing queuing data size, and the target server is selected in the priority order of the queuing rate, the file reading concurrency and the remaining disk space, it can accurately reflect the actual processing capacity and load status of the server, and avoid allocating files to servers with poor performance or heavy load. Compressing text type files can reduce storage space occupancy and improve storage efficiency. The method of selecting the first-level storage directory by the file name hash value and selecting the second-level storage directory by the polling count value makes the file distribution more uniform, reduces the number of files under a single directory, and improves file retrieval efficiency. Aggregate and store multiple small files in the same target aggregation file to reduce the file system load. By recording the compression identifier in the file format field, it is possible to correctly identify whether the file needs to be decompressed when reading. Generate a file name containing complete storage location information so that the specific storage location can be quickly located when accessing the file later, reducing the retrieval overhead.
[0024] 2. The present application provides a distributed high-concurrency aggregate storage method for small files. By detecting the size of the target aggregate file and defragmenting it when it reaches the preset capacity threshold, invalid data blocks can be cleaned up in time to avoid wasting storage space. Valid data blocks are stored continuously again to reduce file fragmentation and improve the utilization efficiency of storage space. The offset information in the file name is updated to ensure that the file content can still be correctly accessed after the data block position changes. This dynamic maintenance mechanism enables the aggregate file to always maintain a good storage state, avoiding the problem of storage space fragmentation caused by the increase in file deletion and update operations. Through regular defragmentation, a high storage space utilization rate can be maintained. At the same time, due to the continuous storage of data blocks, file reading performance can also be improved.
[0025] 3. The present application provides a distributed high-concurrency aggregate storage method for small files. When receiving a file access request, the system can directly locate the target data block information through the index table, and then quickly read the required data from the aggregate file. This indexing mechanism avoids the process of searching one by one in a large number of files, significantly reducing the time complexity of file retrieval. At the same time, the system decompresses the read data to achieve the unity of compressed storage and fast access. By establishing a file name index table, the system can ensure fast retrieval and access of data while maintaining a high compression ratio, effectively balancing the relationship between storage efficiency and access efficiency, and improving the overall data processing capability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flow chart of a small file distributed high-concurrency aggregate storage method in an embodiment of the present application.
[0027] Figure 2 It is another process schematic diagram of a method for distributed high-concurrency aggregated storage of small files in an embodiment of the present application.
[0028] Figure 3 It is a schematic structural diagram of an entity device of a distributed high-concurrency aggregated storage system for small files provided in an embodiment of the present application. Specific embodiments
[0029] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any and all possible combinations including one or more of the listed items.
[0030] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0031] The following uses an embodiment and combines Figure 1 to describe a method for distributed high-concurrency aggregated storage of small files in an embodiment of the present application: Please refer to Figure 1 which is a process schematic diagram of a method for distributed high-concurrency aggregated storage of small files in an embodiment of the present application.
[0032] S101. Obtain the file type and file size of the file to be stored uploaded by the client, and each file server collects performance index data at preset time intervals and sends it to the resource coordination server; The system obtains the file type and file size of the file to be stored uploaded by the client, and each file server collects performance index data at preset time intervals and sends it to the resource coordination server. The performance index data includes the real-time CPU processing speed calculated based on the initial reference value of the CPU main frequency combined with the historical data processing volume, the file processing queue data size, the remaining disk space, and the file reading concurrency.
[0033] The real-time CPU processing speed calculated based on the initial reference value of the CPU main frequency and the historical data processing volume, specifically including: obtaining the total historical data processing volume and the processing completion time within a preset time window; dividing the total historical data processing volume by the processing completion time to obtain the actual processing speed; performing weighted calculation on the initial reference value of the CPU main frequency and the actual processing speed according to a preset weighting coefficient to obtain the real-time CPU processing speed.
[0034] In this step, the system first obtains the basic information of the file to be stored from the client, including the file type and the file size. The file type can be various formats such as text, image, audio and video, and the file size reflects the data magnitude of the file. The purpose of obtaining this information is to provide a basis for subsequent storage optimization and resource scheduling. In addition to passively receiving file uploads, the system can also actively scan a specified directory or device to automatically discover and collect files that need to be stored. At the same time, each file server node will regularly collect its own performance metric data and report it to the central resource coordination server. The performance metrics include CPU processing speed, task queuing situation, disk space, concurrent access volume, etc., and these metrics reflect the real-time load and processing capacity of the server from different aspects.
[0035] In specific implementation, the client first sends the meta-information (file name, type, size, etc.) of the file to be stored to the access layer component of the system in the form of an API request. After the access layer authenticates and checks the legality of the request, it parses out the meta-information and writes it into the task queue for subsequent processing. At the same time, a metric collection module is deployed on each file server, which collects various performance metrics such as CPU, memory, disk, and network at a preset time interval (such as every 5 minutes) and generates a structured metric data packet. The data packet is sent to the resource coordination server through a message queue or an HTTP request. The coordination server persists the received performance metrics into a time series database for subsequent trend analysis and anomaly detection.
[0036] S102. Receive the queuing rate of each file server calculated by the resource coordination server; The system receives the queuing rate of each file server calculated by the resource coordination server. The queuing rate is the ratio of the file processing queuing data size to the real-time CPU processing speed.
[0037] In this step, the file storage scheduling module of the system receives the real-time queuing rate data of each file server node from the resource coordination server. The queuing rate refers to the ratio of the amount of task data queuing for processing on the current server to the data processing speed of the server. It reflects the busyness and load pressure of the server and is an important reference indicator for task scheduling and load balancing. The resource coordination server calculates their respective queuing rates in real time by aggregating the performance metrics reported by each node. In addition to the queuing rate, the coordination server can also calculate some other scheduling-related metrics, such as task completion rate, average response latency, etc., to support more comprehensive and refined task scheduling strategies.
[0038] In specific implementation, the resource coordination server polls the performance metrics of each file server regularly (such as every minute), focusing on obtaining the CPU processing speed and the amount of task queuing data. Then, the coordination server calculates a real-time queuing rate value for each file server according to the preset queuing rate calculation formula. The simplest formula is to divide the queuing data volume by the CPU processing speed, indicating how long it takes to process the currently queued data volume. Considering the cold start factor when the system is just started, a smoothing coefficient can be introduced into the calculation of the queuing rate to give a buffer period to newly added or nodes with suddenly changed processing speeds, avoiding drastic fluctuations in scheduling. In addition, due to the different complexities of file processing tasks, simply dividing the data volume by the processing speed may not be accurate. Therefore, a correction coefficient for task complexity can also be considered in the calculation of the queuing rate, assigning different weights to different types and sizes of files to make the queuing rate closer to the actual processing time of the tasks. The resource coordination server encapsulates the calculated queuing rate into a structured message and pushes it to the file storage scheduling module through a long connection or a message queue to trigger a task scheduling.
[0039] S103. Sort according to the queuing rate, file read concurrency, and remaining disk space in the priority order of descending queuing rate, descending file read concurrency, and ascending remaining disk space, and select the file server with the best sorting result as the target server; This step sorts the file servers according to multiple metrics and selects the optimal server as the storage target. The system comprehensively considers three factors: queuing rate, file read concurrency, and remaining disk space, and sorts them in a specific priority order. The lower the queuing rate, the lighter the current load of the server and the stronger its ability to handle new tasks; the lower the file read concurrency, the less the current IO pressure on the server and the more idle resources; the larger the remaining disk space, the more abundant the storage capacity of the server. The system arranges these three metrics in descending, descending, and ascending order of priority respectively, conducts a comprehensive evaluation of all candidate servers, and the server with the highest score will be selected as the target server for this task. In addition to the above three metrics, the system can also introduce other measurement dimensions such as the memory occupancy rate, network latency, and IO load of the server to build a more comprehensive server evaluation system.
[0040] In specific implementation, the system deploys a weight calculation module in the resource coordination server, which is specifically responsible for calculating the weight coefficients and comprehensive scores of various metrics. The weight calculation module obtains the raw data such as the real-time queuing rate, file read concurrency, and disk usage rate of the candidate servers from the scheduling engine, and then converts each metric into a normalized score between 0 and 100 according to the preset threshold range and calculation formula. Among them, the queuing rate and concurrency adopt the interval division method, and the smaller the value, the higher the score; the disk space adopts the inverse proportional function, and the larger the space, the higher the score. The module also supports customizing and adding new metrics, and flexibly setting the weight coefficients of each metric through the configuration file. Finally, the system adds up the weighted scores of all metrics to obtain the comprehensive performance score of each server, and generates a server ranking list according to the score. The server at the top of the list is the target server with the best performance and will be preferentially selected to execute this storage task.
[0041] S104. When the file type of the file to be stored is a text type, compress the file to be stored to obtain a compressed file and a compression flag; when the file type of the file to be stored is a non-text type, use the file to be stored and an uncompressed flag as the file to be written; When the file type of the file to be stored is a text type, the system performs compression processing on the file to be stored to obtain a compressed file and a compression identifier; when the file type of the file to be stored is a non-text type, the file to be stored and an uncompressed identifier are used as the file to be written. Among them, the system performs compression processing on the file to be stored to obtain a compressed file and a compression identifier, specifically including: obtaining the file size of the file to be stored, and when the file size is greater than the first preset threshold, dividing the file to be stored into multiple data blocks; performing data feature analysis on each data block to obtain the repetition degree, entropy value, and data distribution characteristics of each data block; selecting a corresponding candidate compression algorithm set from a preset compression algorithm library according to the data characteristics; respectively using the compression algorithms in the candidate compression algorithm set to perform compression tests on the data blocks to obtain the compression ratio and compression time of each compression algorithm; based on a preset weight coefficient, calculating the comprehensive score of each compression algorithm according to the compression ratio and compression time, and selecting the compression algorithm with the highest score as the target compression algorithm; using the target compression algorithm to perform compression processing on the data blocks to obtain the compressed data blocks; using the algorithm identifier of the target compression algorithm as the compression identifier, and combining the compressed data blocks to obtain a compressed file.
[0042] This step determines whether to perform compression processing on the file to be stored according to the file type. If the file is of text type (such as TXT, XML, JSON, etc.), the system will compress it to reduce the storage space occupancy; if it is of non-text type (such as pictures, audio and video, etc.), the original file will be directly used as the target file to be written. For text files that need to be compressed, the system will also generate a compression identifier to store metadata such as the file compression algorithm and parameters. The preprocessing stage can be selectively performed according to the file type, or compression can be attempted on all files, and whether to adopt the compressed version is determined according to the compression effect. In addition to the file type, the system can also dynamically determine whether to perform compression according to other factors such as the file size and content characteristics.
[0043] In specific implementation, the system first determines whether the file to be stored belongs to the text type by using features such as the file name suffix and magic number. For files determined to be of the text type, the system further checks whether its size exceeds a preset threshold (such as 1MB). If the file is large, it is first split into multiple data blocks, and each data block is compressed independently. The system will perform data feature analysis on each data block, extract statistical information such as the repetition degree and entropy value of each block, and then use classification algorithms such as decision trees to automatically select multiple candidate compression algorithms with higher matching degrees (such as Gzip, Bzip2, LZMA, etc.). Next, the system uses these candidate algorithms to perform trial compression on the data blocks respectively, records the compression ratio and time consumption of each algorithm, and then comprehensively calculates the performance score. The compression ratio represents the ratio of the size of the compressed data to the original size. The smaller the compression ratio, the more storage space is saved; the time consumption represents the execution time of the algorithm, and the shorter the time consumption, the faster the compression speed. The system uses a weighted scoring method to assign different weights to the two indicators of compression ratio and time consumption, and comprehensively evaluates the performance advantages and disadvantages of each candidate algorithm. Finally, the system selects the algorithm with the highest performance score as the final compression scheme, and generates a compression identifier recording the algorithm type and parameters, which is written into the storage system together with the compressed data. For non-text type or small-volume files, the system directly writes the original file and an identifier indicating the uncompressed state into the storage node.
[0044] S105. Obtain the current polling count value in the target server, and perform a hash operation on the file name of the file to be stored to obtain a first-level directory hash value; This step first obtains a polling count value in the selected target server, which is used to select the secondary storage directory later. The polling count value can be a cyclic incrementing integer, which is used to evenly distribute files among multiple secondary directories to avoid too many files in a single directory. At the same time, the system also performs a hash operation on the file name of the file to be stored to obtain a first-level directory hash value, which is used to determine the first-level storage directory of the file. Using the hash value as the naming method for the first-level directory can make the file distribution more uniform and improve the concurrent access performance. In addition to the hash operation, the system can also select the first-level directory according to other factors such as file type and business attributes.
[0045] During specific implementation, the system first obtains a global polling count value from the target server through methods such as the RPC interface or reading shared memory. Then, the system uses a predefined hash function (such as MD5, SHA-1, etc.) to calculate the file name of the file to be stored, obtaining a hash value of a fixed length. Next, the system performs a modulo operation on the hash value to obtain a remainder less than the number of first-level directories, which serves as the first-level directory hash value of the file. For example, if the system predefines 100 first-level directories numbered from 0 to 99, the remainder obtained by taking the modulo of the hash value by 100 is the first-level directory number. In this way, the system can quickly and evenly allocate storage locations for each file according to the file name, avoiding the problem of unbalanced directory distribution.
[0046] S106. Select a first-level storage directory according to the first-level directory hash value, select a second-level storage directory according to the polling count value, and increment the polling count value cyclically; In this step, according to the first-level directory hash value and polling count value calculated previously, the specific storage path of the file is selected. The system first uses the first-level directory hash value as an index to select the corresponding directory in the predefined first-level directory list as the top-level path for file storage. Then, the system selects the second-level directory with the corresponding number under this first-level directory according to the current polling count value as the final storage location of the file. After completing the selection of the second-level directory, the system increments the polling count value by one and takes the modulo of the total number of second-level directories to ensure that each second-level directory can be cyclically used during the next selection. Through the division of two levels of directories, the system can horizontally expand the storage space while effectively controlling the number of files in a single directory, avoiding the problem of decreased search efficiency caused by an overly large directory.
[0047] During specific implementation, the system maintains a global polling counter with an initial value of 0. When a new file needs to be stored, the system reads the current polling count value, assumed to be n. If the first-level directory hash value is x and the number of second-level directories is m, the storage path of the file can be expressed as " / primaryDir_x / subDir_(n%m)". Here, primaryDir_x represents the first-level directory numbered x, and subDir_(n%m) represents the second-level directory numbered n%m under the first-level directory. % is the modulo operator, ensuring that the second-level directory number cycles within the range of [0, m - 1]. After completing the path calculation, the system increments the polling counter by one and writes it back to storage for use in the next file storage. At the same time, if the polling count reaches the preset maximum value (such as 2^32 - 1), the system resets it to 0 to avoid integer overflow.
[0048] S107. Create a target aggregation file in the secondary storage directory, generate a 2-byte file format field according to the compression flag or uncompressed flag, and generate a 4-byte file length field according to the size of the file to be written; In this step, the system creates a new aggregation file under the selected secondary storage directory to store multiple small files to be written. The aggregation file can significantly reduce the metadata overhead in the file system and improve the storage and access efficiency of a large number of small files. At the same time, to support the parsing and management of the aggregation file, the system also needs to generate corresponding format identification fields and length fields when writing each small file. Among them, the file format field occupies 2 bytes and is used to identify whether the current file is compressed; the file length field occupies 4 bytes and is used to record the actual size of the current file. Through these two fields, the system can accurately locate the start and end positions of each small file when reading the aggregation file, realizing efficient random access.
[0049] Specifically, the system first generates a unique aggregation file name according to information such as the current date, time, and server ID. Then, it creates this aggregation file in the specified secondary directory and opens a file handle to prepare for writing data. Next, the system checks the compression flag of the file to be written. If it is compressed, the file format field is set to 0x0001; if it is uncompressed, the file format field is set to 0x0000. Using a 2-byte short integer to store the file format is to be compatible with other systems and protocols, and at the same time reserve an extension space for new formats that may appear in the future. After setting the file format field, the system then generates a 4-byte unsigned integer as the file length field according to the actual size of the file to be written. The reason for using a 4-byte integer is that the upper limit of the size of a single file in most file systems is 4GB (2^32 bytes), and using a 4-byte length field can cover the storage requirements of most small files. For the extremely few large files exceeding 4GB, the system can cut them before writing to ensure that the size of each sliced file does not exceed 4GB.
[0050] S108. Write the file format field, the file length field, and the file to be written into the target aggregation file in sequence, and record the starting offset of the writing; This step writes the previously generated file format field, file length field, and the file content to be written into the target aggregated file. When writing, it is necessary to strictly follow the order of the format field, length field, and file content to ensure that the boundaries of each small file can be accurately parsed during subsequent reading. At the same time, to support the fast positioning and random access of small files, after the writing is completed, the system also needs to record the starting offset of each small file in the aggregated file. The offset can be the number of bytes relative to the starting position of the aggregated file or the sequence number of the small file in the aggregated file. Associating the offset information with the metadata of the small file (such as file name, creation time, etc.) for storage can significantly speed up the retrieval of small files.
[0051] In specific implementation, the system first writes the file format field and file length field to the current position of the aggregated file. Since the lengths of these two fields are fixed (2 bytes and 4 bytes respectively), the underlying write interface of the file system can be directly used during the writing process without additional buffering or format conversion. Next, the system continuously writes the content of the file to be written into the aggregated file. To improve the writing efficiency, the system can set a write buffer for the aggregated file, first accumulate the content of multiple small files in the buffer, and then write it to the disk at one time to reduce the number of I / Os. After the writing is completed, the system needs to obtain the current position of the file pointer, subtract the position before writing, and get the number of bytes written this time. The number of bytes corresponds one-to-one with the starting offset of the current small file in the aggregated file. The system records this pair of offset information in the memory index or index file for subsequent reading and retrieval. To further improve the retrieval efficiency, the system can also compress and merge the index information, and periodically persist the index in memory to a disk file to avoid loss after restart.
[0052] S109. Generate a file name containing storage location information based on the server identifier of the target server, the first-level directory hash value, the polling count value, the target aggregated file identifier, and the starting offset, and return the file name to the client.
[0053] This step generates a unique file name containing the complete storage location information based on the multiple pieces of information obtained previously, as the identifier of the file to be stored in the distributed system, and returns the file name to the client for the client to use when accessing and operating on the file subsequently. The storage location information contained in the file name can help the client quickly and accurately locate the file without having to query the file index, thus improving the file access efficiency. The system can also add other auxiliary information, such as file version number, checksum, etc., to the file name to support more functional requirements.
[0054] In the above embodiment, by calculating the real-time CPU processing speed based on the initial reference value of the CPU main frequency and the historical data processing volume, the queuing rate is obtained in combination with the file processing queuing data size, and the target server is selected according to the priority order of the queuing rate, the file reading concurrency and the remaining disk space, the actual processing capacity and load status of the server can be accurately reflected, and the files can be avoided from being allocated to servers with poor performance or heavy load. Compressing text type files can reduce storage space occupancy and improve storage efficiency. The method of selecting the first-level storage directory by the file name hash value and selecting the second-level storage directory by the polling count value makes the file distribution more uniform, reduces the number of files under a single directory, and improves file retrieval efficiency. Aggregate and store multiple small files in the same target aggregate file to reduce the file system load. By recording the compression identifier in the file format field, it is possible to correctly identify whether the file needs to be decompressed when reading. Generate a file name containing complete storage location information so that the specific storage location can be quickly located when accessing the file later, reducing the retrieval overhead.
[0055] In the above embodiment, by selecting a suitable target server and aggregating and storing small files, efficient distributed storage of files is achieved. However, during the long-term operation of the system, due to frequent file updates and deletions, a large number of storage fragments may be generated in the aggregated files, affecting the storage efficiency and access performance of the system. Figure 2 , another small file distributed high-concurrency aggregate storage method in the embodiment of the present application is described: See also Figure 2 , is another flow chart of a small file distributed high-concurrency aggregate storage method in an embodiment of the present application.
[0056] S201, detecting whether the current file size of the target aggregate file reaches a preset capacity threshold; The threshold is usually determined according to the system's storage policy and performance requirements, for example, it can be set to 512MB, 1GB, etc. When the size of the aggregated file exceeds the threshold, it is necessary to trigger subsequent defragmentation operations to improve storage space utilization and file access performance. In addition to the fixed capacity threshold, the system can also dynamically adjust the threshold size according to the actual storage space usage to balance storage efficiency and defragmentation overhead.
[0057] In specific implementation, the system first obtains the metadata information of the target aggregated file, including file name, creation time, last modification time, file size, etc. These metadata can be obtained from the index structure of the file system or a dedicated metadata storage. Then, the system compares the obtained file size with a preset capacity threshold to determine whether it reaches or exceeds the threshold. If it does not reach the threshold, the detection process is terminated and normal file read and write operations continue; if it reaches or exceeds the threshold, the defragmentation process is triggered and the next operation is entered. In an actual system, to reduce the detection overhead, a periodic polling or asynchronous notification method can be adopted, and a detection interval time can be set to avoid performing detection operations too frequently. In addition, the system can set different capacity thresholds for different aggregated files to adapt to the storage characteristics of different types of files.
[0058] S202. When the current file size reaches the preset capacity threshold, defragment the data in the target aggregated file; When the current file size reaches the preset capacity threshold, the system defragments the data in the target aggregated file, specifically including: constructing a file mapping table to record the starting position, length, and valid status of each data block in the target aggregated file; traversing the file mapping table in the storage order of the data blocks; sequentially migrating the data blocks with valid status to the starting position of the file to obtain continuously stored valid data; and updating the position information of each valid data block in the file mapping table.
[0059] After detecting that the size of the target aggregated file reaches the preset capacity threshold, this step defragments the storage space in the file. The aggregated file adopts an append-write method to continuously store the data of multiple small files. However, during the long-term operation of the system, due to file update and deletion operations, storage holes will be generated in the aggregated file, forming internal fragmentation. The main purpose of defragmentation is to reorganize the valid data into a continuous storage space, improve the space utilization rate, and accelerate the file reading speed. The defragmentation process can be performed online, offline, or an incremental defragmentation strategy can be adopted to balance storage performance and defragmentation efficiency.
[0060] During specific implementation, the system first scans the target aggregation file to identify and record the valid data blocks and invalid data blocks in the file. Valid data blocks refer to the file data that is currently still in use and has not been deleted or updated; invalid data blocks refer to the file data that has been deleted or overwritten by new version data. The identification method can be based on the flag bits in the file header or by comparing the records in the file index table. After the scan is completed, the system obtains a complete file mapping table, which records information such as the starting position, data length, and valid status of each data block. Next, the system traverses the records in the mapping table in the order of the data blocks in the file. For data blocks with a valid status, the system copies or moves them to the starting position of the file and updates the position information of the corresponding data blocks in the mapping table; for data blocks with an invalid status, the system directly skips them without any processing. After traversing all data blocks, the system reorganizes the valid data into a continuous storage space, discards the invalid data, and tidies up the fragmentation. Finally, the system updates the metadata information of the file, such as the file size, modification time, etc., to complete the entire defragmentation process.
[0061] S203. Delete the data blocks marked as invalid in the target aggregation file and store the valid data blocks continuously again; After this step completes the defragmentation, according to the defragmentation result, delete the invalid data blocks in the aggregation file and reorganize the valid data blocks into a continuous storage space. By deleting the invalid data, the occupied storage resources can be released, improving the storage space utilization rate; by reorganizing the valid data, the internal layout of the file can be optimized, reducing external fragmentation and accelerating the file reading and writing speed. The deletion and reorganization operations of the data blocks can directly modify the file in place or be completed in another temporary file and then replace the original file.
[0062] S204. Update the offset information in the file name corresponding to the valid data blocks.
[0063] After reorganizing the valid data blocks, this step updates the offset information in the file name corresponding to each valid data. Since the file name adopts an encoding method that includes storage location information, after the position of the data block in the aggregation file changes, it is necessary to synchronously update the offset field in the file name to ensure that subsequent file access can correctly locate the data block. The update of the offset information can be completed synchronously, asynchronously modified, or selectively update the hot files with a higher access frequency.
[0064] In specific implementation, the system first traverses the file mapping table obtained in the previous step to find all data blocks in the valid state. For each valid data block, the system reads its starting position in the reorganized aggregated file from the mapping table and calculates the offset relative to the file header. Then, based on information such as the length of the data block and the ID of the file it belongs to, the system searches for the corresponding file name record in the distributed file index table. Here, the index table can be a Key-Value type metadata storage, with the file ID as the Key and the file name and other metadata as the Value. After finding the file name record, the system fills the new offset into the corresponding field of the file name string to complete the file name update for one data block. Repeat this process until the file names of all valid data blocks are updated. The updated file names match the new positions of the data blocks in the aggregated file, and the data access logic can continue to work without modifying the upper-layer business code. The system can also choose to update the file names in batches. First, calculate the offsets of all data blocks, and then submit them to the metadata storage at one time to reduce the number of update operations and improve the processing efficiency.
[0065] In the above embodiment, by detecting the size of the target aggregated file and performing defragmentation when reaching the preset capacity threshold, invalid data blocks can be cleaned up in time to avoid waste of storage space. Storing the valid data blocks continuously reduces file fragmentation and improves the utilization efficiency of storage space. Updating the offset information in the file name ensures that the file content can still be accessed correctly after the position of the data block changes. This dynamic maintenance mechanism enables the aggregated file to always maintain a good storage state and avoids the problem of storage space fragmentation caused by the increase in file deletion and update operations. Through regular defragmentation, a high storage space utilization rate can be maintained. At the same time, due to the continuous storage of data blocks, the file reading performance can also be improved.
[0066] Further, in another embodiment, when updating the position information of each valid data block in the file mapping table, the method further includes: establishing a file name index table to store the mapping relationship between the file name and the corresponding data block. Among them, establishing the file name index table specifically includes: parsing the file name into a server identifier, directory information, file identifier, and offset; locating the target storage location based on the server identifier and directory information; establishing a mapping relationship with the target data block according to the file identifier and offset to obtain the file name index table; when detecting a file access request, obtaining the target data block information from the file name index table according to the file access request; reading the data from the target aggregated file based on the target data block information; and performing decompression processing on the read data.
[0067] First, the system needs to build a file name index table to store the mapping relationship between file names and corresponding data blocks. Since the file names use a special encoding method, which contains multiple fields such as server identification, directory information, file identification, and offset, the file names need to be parsed first to extract the values of each field. The parsing process can use methods such as regular expressions or string splitting to match and extract according to predefined format rules. Then, based on the parsed server identification and directory information, the system locates the target storage location where the file is located, such as a specific physical server and disk partition. Next, according to the file identification and offset fields, the system searches for the corresponding data block in the target storage location and establishes a mapping relationship from the file name to the data block, forming an entry in the file name index table. Repeat this process of parsing, locating, and mapping until all file names are processed to obtain a complete file name index table.
[0068] After the file name index table is established, it can significantly accelerate file access operations. When the system detects a file access request initiated by the upper-layer application, it first obtains the file name of the target file from the request, and then uses this file name as the Key to search in the file name index table. If a matching entry is found, the storage location information of the target data block can be directly obtained, including the aggregated file it is in, the offset, the data length, etc. In this way, the system can quickly locate the target position in the aggregated file according to the data block information and read the data block content of the corresponding length, avoiding the overhead of full-text scanning and traversal.
[0069] After reading the target data block, since the data may be compressed, corresponding decompression operations need to be performed to restore the original file content. The specific decompression method can be determined according to the compression flag bit in the file name. If the flag indicates a compressed file, the corresponding decompression algorithm is called for processing; if the flag indicates an uncompressed file, the read data is directly returned. The choice of decompression algorithm can combine the characteristics of the data block and the system's resource situation and adopt an efficient implementation method, such as hardware acceleration, parallel computing, etc. The decompressed file content can be returned to the upper-layer application to complete the entire file access process.
[0070] In the above embodiments, when a file access request is received, the system can directly locate the target data block information through the index table, and then quickly read the required data from the aggregated file. This indexing mechanism avoids the process of searching one by one in a large number of files, significantly reducing the time complexity of file retrieval. At the same time, the system decompresses the read data, achieving the unity of compressed storage and fast access. By establishing the file name index table, the system can still ensure the fast retrieval and access of data while maintaining a high compression ratio, effectively balancing the relationship between storage efficiency and access efficiency, and improving the overall data processing ability of the system.
[0071] The system in the embodiments of the present invention application will be described below from the perspective of hardware processing. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of a small file distributed high-concurrency aggregated storage system provided by the embodiments of the present application.
[0072] It should be noted that Figure 3 the structure of the system shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present invention.
[0073] As Figure 3 shown, the system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 302 or the program loaded from the storage section 308 into the Random Access Memory (RAM) 303, such as executing the method in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The Input / Output (I / O) interface 305 is also connected to the bus 304.
[0074] The following components are connected to the I / O interface 305: an input section 306 including a camera, an infrared sensor, etc.; an output section 307 including a liquid crystal display (LCD) and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card and a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.
[0075] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, various functions defined in the present invention are executed.
[0076] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above.
[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0078] As another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or may exist alone without being assembled into the system. The above storage medium carries one or more computer programs, and when the one or more computer programs are executed by a processor of a system, the system implements the method provided in the above embodiments.
[0079] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0080] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted as "if determining...", "in response to determining...", "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0081] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.
[0082] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM, random access memory (RAM), magnetic disks, or optical discs.
Claims
1. A distributed high-concurrency aggregation storage method for small files, characterized in that Including: Obtain the file type and file size of the file to be stored uploaded by the client; Each file server collects performance metric data at preset time intervals and sends it to the resource coordination server. The performance metric data includes the real-time CPU processing speed calculated based on the initial reference value of the CPU main frequency combined with the historical data processing volume, the file processing queue data size, the remaining disk space, and the file reading concurrency; Receive the queuing rate of each file server calculated by the resource coordination server. The queuing rate is the ratio of the file processing queue data size to the real-time CPU processing speed; Sort according to the queuing rate, the file reading concurrency, and the remaining disk space in the priority order of descending queuing rate, descending file reading concurrency, and ascending remaining disk space, and select the file server with the best sorting result as the target server; When the file type of the file to be stored is a text type, perform compression processing on the file to be stored to obtain a compressed file and a compression identifier; When the file type of the file to be stored is a non-text type, use the file to be stored and an uncompressed identifier as the file to be written; Obtain the current polling count value in the target server, and perform a hash operation on the file name of the file to be stored to obtain a first-level directory hash value; Select a first-level storage directory according to the first-level directory hash value, select a second-level storage directory according to the polling count value, and increment the polling count value cyclically; Create a target aggregation file in the second-level storage directory, generate a 2-byte file format field according to the compression identifier or the uncompressed identifier, and generate a 4-byte file length field according to the size of the file to be written; Write the file format field, the file length field, and the file to be written into the target aggregation file in sequence, and record the starting offset of the writing; Generate a file name containing storage location information according to the server identifier, the first-level directory hash value, the polling count value, the target aggregation file identifier, and the starting offset of the target server, and return the file name to the client.
2. The method according to claim 1, wherein The real-time CPU processing speed calculated based on the initial reference value of the CPU main frequency combined with the historical data processing volume specifically includes: Obtain the total historical data processing volume and the processing completion time within a preset time window; Divide the total historical data processing volume by the processing completion time to obtain the actual processing speed; Perform weighted calculation on the initial reference value of the CPU main frequency and the actual processing speed according to a preset weighting coefficient to obtain the real-time CPU processing speed.
3. The method according to claim 1, wherein The performing compression processing on the file to be stored to obtain a compressed file and a compression identifier specifically includes: Obtain the file size of the file to be stored. When the file size is greater than a first preset threshold, divide the file to be stored into multiple data blocks; Perform data feature analysis on each data block to obtain the repeatability, entropy value, and data distribution characteristics of each data block; Select a corresponding candidate compression algorithm set from a preset compression algorithm library according to the data characteristics; Compress and test the data block using the compression algorithms in the candidate compression algorithm set respectively, and obtain the compression ratio and compression time of each compression algorithm; Based on the preset weight coefficients, calculate the comprehensive score of each compression algorithm according to the compression ratio and the compression time, and select the compression algorithm with the highest score as the target compression algorithm; Use the target compression algorithm to compress the data block to obtain a compressed data block; Use the algorithm identifier of the target compression algorithm as the compression identifier, and combine the compressed data blocks to obtain a compressed file.
4. The method according to claim 1, wherein After the starting offset of the record writing, the method further includes: Detect whether the current file size of the target aggregation file reaches a preset capacity threshold; When the current file size reaches the preset capacity threshold, defragment the data in the target aggregation file; Delete the data blocks marked as invalid in the target aggregation file, and store the valid data blocks continuously again; Update the offset information in the file name corresponding to the valid data block.
5. The method according to claim 4, characterized in that The defragmenting the data in the target aggregation file specifically includes: Construct a file mapping table to record the starting position, length, and valid status of each data block in the target aggregation file; Traverse the file mapping table in the storage order of the data blocks; Migrate the data blocks with valid status to the starting position of the file in sequence to obtain continuously stored valid data; Update the position information of each valid data block in the file mapping table.
6. The method according to claim 4, characterized in that In the updating the position information of each valid data block in the file mapping table, the method further includes: Establish a file name index table to store the mapping relationship between the file name and the corresponding data block; When a file access request is detected, obtain the target data block information from the file name index table according to the file access request; Read data from the target aggregation file based on the target data block information; Perform decompression processing on the read data.
7. The method according to claim 6, wherein The establishing the file name index table specifically includes: Parse the file name into a server identifier, directory information, file identifier, and offset; Locate the target storage location based on the server identifier and directory information; Establish a mapping relationship with the target data block according to the file identifier and offset to obtain a file name index table.
8. A distributed high-concurrency aggregated storage system for small files, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the system to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the system, cause the system to execute the method according to any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product runs on the system, cause the system to execute the method according to any one of claims 1-7.
Citation Information
Cited By
Data storage method, distributed storage system, device, medium and product
CN120540608A
Remote file differential transmission and iteration method and device
CN120849355A
Method, device and equipment for quickly loading data file of time sequence database
CN121144386A
Discrete file aggregation method for mass data
CN121681479A
Data acquisition method and system with evolution prediction mechanism, and program product
CN121959032A