Hadoop massive small file reading method based on file merging and heat and aging dual elimination mechanism
By introducing the Redis cache module and the dual elimination mechanism of popularity and timeliness in the Hadoop system, caching and merging massive small files has been solved, and the performance bottleneck of Hadoop system when processing massive small files has been achieved, and efficient file reading and processing has been achieved.
Patent Information
- Application Number
- CN202510221087.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
When Hadoop system processes massive small files, NameNode has high memory pressure and slow response speed, and frequent disk I/O operations by traditional methods lead to poor performance.
Using a method based on file merging and dual elimination mechanism of heat and timeliness, hot small files are cached through the Redis cache module to reduce metadata query of NameNode and disk I/O operations of DataNode.
It significantly improves the reading speed of massive small files, reduces the reading delay, and enhances the performance and adaptability of Hadoop system in the processing scenarios of massive small files.
Smart Images

Figure CN120067070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method for reading a large number of small files in Hadoop based on file merging and a dual elimination mechanism of heat and timeliness. Background Art
[0002] In the big data processing ecosystem, Hadoop has become a key technology platform for storing and processing a large amount of data due to its distributed architecture and strong scalability. However, with the continuous expansion of the data scale and the increasing complexity of data types, the problem of processing a large number of small files has gradually become a bottleneck for improving the performance of the Hadoop system. Taking the e-commerce industry as an example, the number of small files such as product reviews and order details generated every day can reach millions or even tens of millions. When these small files are stored in the Hadoop Distributed File System (HDFS), since each small file needs to occupy a certain amount of memory space in the NameNode to store its metadata (such as file name, file size, storage location, etc.). When the number of small files is huge, the memory pressure on the NameNode increases sharply, resulting in a slowdown in the system response speed and even jamming.
[0003] At the same time, when the traditional Hadoop method reads a large number of small files, frequent disk I / O operations are one of the main reasons for low performance. The data volume of small files is small, and each read requires disk seek, head positioning and other operations. The overhead of these operations is continuously accumulated when reading a large number of small files, making disk I / O the performance bottleneck of the entire system. For example, the time taken to read 100 small files of 10KB far exceeds the time taken to read a large file of 1MB, even though the total data volume of the two is the same. This is because a large file can be read at one time, while small files require multiple disk I / O operations, increasing the waiting time of the system.
[0004] In addition, the NameNode, as the core component of HDFS, is responsible for managing the namespace of the file system and file metadata. A large number of small files means a huge storage requirement for metadata, which causes the memory consumption of the NameNode to increase sharply. At the same time, when processing file read requests, the NameNode needs to frequently query and process this metadata, resulting in an overloaded load and a decrease in processing speed. For example, when millions of small file metadata are stored in the NameNode's memory, it may take several seconds or even longer to find the metadata of a specific small file, seriously affecting the efficiency of file reading. Summary of the Invention
[0005] In order to overcome the problem of low efficiency in reading a large number of small files by the above traditional Hadoop method, the present invention provides a method for reading a large number of small files in Hadoop based on file merging and a dual elimination mechanism of heat and timeliness.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A method for reading a large number of small files in Hadoop based on file merging and a dual elimination mechanism of heat and timeliness, the method is applicable to an HDFS system having a data merging module and a Redis cache module, and the method includes the following steps:
[0008] Receive a small file reading request input by a user, and determine whether the file of the reading request is a small file;
[0009] If not, send a request to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information;
[0010] If so, query the Redis cache module according to the complete path of the small file. If the small file requested to be read is cached in the Redis cache module, directly return the read small file to the user; the Redis cache module caches some hot small files according to the cache update policy, and the hot small files are predicted by a small file access prediction module constructed by a heat calculation formula based on a dual elimination mechanism of hot spots and timeliness;
[0011] If the small file requested to be read is not in the Redis cache module, send a request to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information.
[0012] Preferably, the heat calculation formula based on the dual elimination mechanism of hot spots and timeliness includes the following:
[0013]
[0014] Among them, α represents the attenuation coefficient, 0 < α < 1; countPeriod represents the heat calculation period; visitCount represents the number of times the small file is accessed in the current heat calculation period; heat n-1 represents the historical heat of the small file.
[0015] Preferably, before receiving the small file reading request input by the user, the method further includes: initializing the HDFS system;
[0016] The initialization includes:
[0017] Configuring network parameters and cache parameters, and the cache parameters include cache capacity and cache expiration time;
[0018] Start the cache update policy engine and load the pre-trained small file access prediction model; Initialize relevant parameters and variables, including the historical data storage location and the small file access prediction model update frequency;
[0019] Initialize the Redis cache module, set the initial size of the cache space, and configure the relevant parameters of the cache update policy; The relevant parameters of the cache update policy include the preloading time window and the preloading file type.
[0020] Preferably, the cache update policy includes:
[0021] When executing a read request, only record the number of times a file is accessed within the heat calculation period until the number of requests reaches countPeriod, that is, when the heat calculation period ends, the service process triggers the cache update and replacement operation;
[0022] At the same time, calculate the heat value of all files according to the heat calculation formula, sort according to the heat value, and cache the files with the top-K heat values and their metadata into the Redis cache module using the "insert-on-access" strategy;
[0023] When the cache space of the Redis cache module is full, the file with the smallest heat value will be selected and evicted from the cache according to the heat value of the file, and the files with the top-K heat values will be retained in the cache, thus realizing the dynamic update and optimization of the cache data.
[0024] Furthermore, the cache update policy further includes: using the system idle period or low load period to pre-load the files with the top-K heat values into the Redis cache module in advance.
[0025] Preferably, the data merging method adopted by the data merging module is as follows:
[0026] First, analyze the historical access logs of small files to obtain the access situation of each small file by users;
[0027] Calculate the overall relevance of a certain small file to all files in a merging queue;
[0028] When the overall relevance reaches the preset merging threshold, merge the small file into the merging queue and store it in HDFS by calling the API of MapFile provided by Hadoop; The file structure of MapFile consists of two parts: index and data, which are used to store the index file and data respectively.
[0029] Furthermore, the calculation of the overall relevance of a certain small file to all files in a merging queue includes:
[0030] Correlation(f, fi) = {(f, fi)|P(fi|f) ∈ (th1, th2) * P(ffi) ∈ (th3, th4)} (1)
[0031]
[0032] Among them, in Equation (1), Correlation(f, fi) represents the degree of association between file f and file fi; (f, fi) represents the relational binary group composed of files f and fi; P(fi|f) represents the probability that file fi is accessed within a period of time after file f is accessed; P(ffi) represents the probability that the user accesses both f and fi within a period of time; th represents different thresholds, and their value ranges are all within the interval (0, 1).
[0033] In Equation (2), Correlation(f, q) represents the overall correlation degree between the current small file f and all files in a certain merging queue q; n represents the number of merging queues.
[0034] Preferably, the method further includes a file writing process, and the file writing process includes:
[0035] Receiving a file writing request sent by the user, and determining whether the written file is a small file;
[0036] If it is a small file, then enter the data merging module for data merging, and transfer the merged small file to HDFS for storage;
[0037] If it is not a small file, then directly enter the HDFS system for storage operation.
[0038] Preferably, the method further includes a file deletion process, and the file deletion process includes:
[0039] Receiving a file deletion request sent by the user; determining whether the deleted file is a small file;
[0040] If it is not, then directly enter the HDFS system for deletion operation;
[0041] If it is, then query the Redis cache module according to the complete file path; if the file to be deleted is in the Redis cache module, then perform a deletion operation on the small file cached in the Redis cache module;
[0042] If the file to be deleted is not in the Redis cache module, then send a deletion request to the NameNode in the Hadoop cluster; the NameNode matches and interacts with the DataNode to find the corresponding small file, sends a deletion request for deletion operation; and returns the information of the deleted small file to the user.
[0043] Preferably, the method further includes:
[0044] When an update, deletion, or addition operation occurs to a small file in Hadoop, obtain corresponding information through a message queue or an event listening mechanism;
[0045] The Redis cache module synchronously updates, deletes, or adds the small files cached in the Redis cache module according to the received information.
[0046] Compared with the prior art, the beneficial effects of the present invention are:
[0047] By optimizing the distributed system HDFS, combining the small file cache update strategy, and introducing a small file access prediction module in the Redis cache module, the present invention aims to significantly improve the reading speed of a large number of small files, greatly reduce the reading latency, comprehensively enhance the performance and adaptability of the Hadoop system in the scenario of processing a large number of small files, and meet the urgent needs of various industries for real-time big data processing.
[0048] The present invention delves into the core field of big data storage and reading technologies, and focuses on innovatively breaking through the performance dilemmas faced by the Hadoop platform in processing the reading of a large number of small files. In the current era background of accelerating digital transformation, the data volume of various industries has shown an explosive growth. Among them, a large number of small file data widely exist in many business scenarios such as e-commerce transaction records, social platform user dynamics, and financial institution transaction details. As the mainstream platform for big data storage and processing, the reading performance of Hadoop for a large number of small files directly relates to the data value mining efficiency and the smoothness of business operations. The results of the present invention are expected to provide efficient and reliable technical support for the above industries and related fields, and promote the in-depth development of big data applications. Description of the Drawings
[0049] Figure 1 is a flowchart of the steps of the method for reading a large number of small files in Hadoop based on the file merging and dual elimination mechanism of heat and timeliness of the present invention.
[0050] Figure 2 is a flowchart of the cache management process of the present invention.
[0051] Figure 3 is a flowchart of the file merging process of the present invention.
[0052] Figure 4 is a flowchart of the file writing process of the present invention.
[0053] Figure 5 is a flowchart of the file deletion process of the present invention. Detailed Embodiments
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. The following will describe the present invention in detail with reference to the accompanying drawings and specific embodiments.
[0055] It should be understood that when used in this specification, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0056] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0057] It should be further understood that the term "and / or" used in this specification of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0058] Embodiment 1
[0059] As Figure 1 shown, a method for reading a large number of small files in Hadoop based on a file merging and dual elimination mechanism of heat and timeliness, the method is applicable to an HDFS system having a data merging module and a Redis cache module, and the method includes the following steps:
[0060] Receive a small file reading request input by a user, and determine whether the file of the reading request is a small file;
[0061] If not, send a request to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information;
[0062] If so, query the Redis cache module according to the complete path of the small file. If the small file requested to be read is cached in the Redis cache module, directly return the read small file to the user; the Redis cache module caches some hot small files according to the cache update policy, and the hot small files are predicted by a small file access prediction module constructed by a heat calculation formula based on a dual elimination mechanism of hot spots and timeliness;
[0063] If the small file requested to be read is not in the Redis cache module, a request is sent to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information.
[0064] In a specific embodiment, to improve the access efficiency of small files in the HDFS system, an intelligent caching mechanism is designed and implemented. Through the combination of the small file access prediction module and the cache update policy, some hot small files are placed in the Redis cache module for caching, which can reduce the time-consuming operations such as finding metadata in the Namenode and reading Blocks in the Datanode, thus effectively improving the access efficiency of small files in HDFS.
[0065] Based on the reasonable assumption that "the data most recently accessed is most likely to be accessed repeatedly in the near future", the present invention designs a unique dual elimination mechanism. By periodically accumulating the number of times a cached file is accessed and converting it into a heat value, accurate screening of hot data is achieved.
[0066] Among them, the heat calculation formula of the dual elimination mechanism based on hotness and timeliness includes the following:
[0067]
[0068] Among them, α represents the attenuation coefficient, 0 < α < 1; countPeriod represents the heat calculation period; visitCount represents the number of times the small file is accessed in the current heat calculation period; heat n-1 represents the historical heat of the small file.
[0069] The heat calculation formula cleverly balances the weights of the current access times and the historical heat when calculating the heat of a file. The larger α is, the greater the weight of the most recent access in the data access heat, and the smaller the influence of the historical access record on the data heat; conversely, the smaller α is, the relatively greater the influence of the historical heat. During the actual operation process, the historical heat of the file will decay at a rate of (1 - α) coefficient in each calculation period; after multiple iterations, the influence of the early accumulated heat on the data heat gradually decreases, so as to ensure that the heat calculation can timely reflect the real-time access heat of the file.
[0070] In a specific embodiment, in addition to predicting by the small file access prediction module constructed by the heat calculation formula of the dual elimination mechanism based on hotness and timeliness, the intelligent caching mechanism further includes a cache update policy.
[0071] The cache update policy, as Figure 2 shown, includes:
[0072] To reduce the system overhead caused by heat calculation, when the caching service process executes a read request, only the number of times a file is accessed within the heat calculation period is recorded, and the cached data is not immediately replaced. Instead, the service process triggers the cache update and replacement operation until the number of requests reaches countPeriod, that is, at the end of the heat calculation period.
[0073] At the same time, the heat values of all files are calculated according to the heat calculation formula, and the files are sorted according to the heat values. The files with the TOP-K highest heat values and their metadata are cached in the Redis cache module using the "insert-on-access" strategy. When determining the TOP-K heat value ranking, the size limit of the cache space is fully considered, and the heat threshold is calculated to ensure the reasonable utilization of the cache space. In the initial stage of the system, due to a large amount of free cache space, the "insert-on-access" strategy is adopted to insert all accessed files into the cache, quickly establishing an initial cache data set.
[0074] When the cache space of the Redis cache module is full, the file with the lowest heat value is selected and evicted from the cache according to the heat value of the file, and the files with the TOP-K highest heat values are retained in the cache, thus realizing the dynamic update and optimization of the cached data. The cache replacement strategy based on the heat value of the present invention begins to take effect, and small files ranked after TOP–K are selected as "victims" and evicted from the cache according to the heat value of the file, and small files with high heat values are retained in the cache, thus realizing the dynamic update and optimization of the cached data.
[0075] In a specific embodiment, the cache update strategy further includes: using the system idle period or low-load period to pre-load the files with the TOP-K highest heat values into the Redis cache module in advance.
[0076] By introducing the cache preloading strategy, the present invention makes full use of the resources during the system idle period or low-load period. Through the analysis of the system access pattern by the small file access prediction module, the hot files to be pre-loaded are determined. For example, in a news information platform, the early morning is the peak period for users to access news small files every day. The system can, according to the historical access data and heat value prediction during the early morning low-load period, pre-load the news small files and their metadata that may be accessed in large quantities on that day into the cache in advance. In this way, during the high-concurrency access in the morning, the data requested by users can be directly obtained from the cache, greatly improving the system response speed and user experience.
[0077] The present invention collects and stores historical access data of small files, covering rich information such as access time and access frequency. Using advanced machine learning algorithms, such as decision trees, neural networks, etc., these historical data are deeply mined and analyzed. Combining the heat calculation formula with a dual elimination mechanism based on heat and timeliness, a highly accurate small file access prediction model is constructed. Taking a social media platform as an example, by analyzing the access frequency and patterns of picture small files and text small files by users at different time periods (such as weekdays, weekends, holidays) and different activities (such as posting dynamics, browsing friends' dynamics, liking and commenting, etc.), a model that can accurately predict the access probability of small files in different scenarios is trained. This model not only considers the impact of time factors on file access, but also combines user behavior characteristics, improving the accuracy and reliability of prediction.
[0078] After constructing the small file access prediction module, as time goes by and new access data accumulates continuously, the data access pattern may change. In order to keep the small file access prediction module always highly accurate, the small file access prediction module is retrained and optimized regularly. When the platform launches new functions (such as short video functions) or major changes occur in the business scenario (such as the adjustment of the user group structure), the input features and training parameters of the small file access prediction module are adjusted in a timely manner. For example, after introducing the short video function, the access data related to short videos is included in the training scope of the small file access prediction module, and the feature weights are adjusted to adapt to the new access pattern, ensuring that the small file access prediction module can continuously and accurately predict the access trend of small files and provide a reliable basis for intelligent prefetching.
[0079] In a specific embodiment, before receiving a small file reading request input by a user, the method further includes: initializing the HDFS system;
[0080] The initialization includes:
[0081] Configuring network parameters and cache parameters, where the cache parameters include cache capacity and cache expiration time.
[0082] The present invention carefully deploys and configures the Redis cache module in the Hadoop cluster to ensure a stable and efficient connection channel is established between it and each node of Hadoop. By reasonably configuring network parameters and cache parameters (such as cache capacity, cache expiration time, etc.), it guarantees that data can be quickly transmitted and interacted between the Hadoop cluster and the Redis cache, providing a solid foundation for subsequent cache operations. The Redis cache module is also used to receive and process write requests initiated by the client. When a write request is received, according to the cache storage strategy, the cache is stored in the Redis cache module for management, which will greatly improve the read requests initiated by the client. At the same time, the Redis cache module is also responsible for managing all the meta-information data in small files. The Redis cache module works in coordination with the Hadoop cluster. The Redis cache module shares the concurrent pressure brought by a large number of files and greatly improves the reading efficiency of hot files.
[0083] Start the cache update strategy engine and load the pre-trained small file access prediction model; initialize relevant parameters and variables, including the storage location of historical data (specify the database table or file path for storing historical access data) and the update frequency of the small file access prediction model (set to once a week or dynamically adjusted according to business requirements). Ensure that the prefetch decision engine can operate accurately to provide strong support for intelligent prefetching.
[0084] Initialize the Redis cache module, set the initial size of the cache space, and configure the relevant parameters of the cache update strategy; the relevant parameters of the cache update strategy include the preloading time window (for example, set the preloading time window from 2 am to 5 am) and the preloading file type (specified according to business requirements, such as for an e-commerce platform, it can specify types such as product pictures and user review files). By initializing the Redis cache module, it prepares for the effective management of the cache and the preloading of hot files.
[0085] In the initial stage of the system in this embodiment, due to sufficient free space in the cache, the "access and insert" strategy is adopted to insert all accessed files and their metadata into the Redis cache module for caching, quickly constructing an initial cache data set. When the Redis cache module cache gradually fills up, a dual elimination mechanism based on popularity and timeliness and a cache preloading strategy are combined to dynamically replace the hot files and their metadata in the cache. After each file access operation, the cache service process records the access times of the file. When the popularity calculation period is reached, the file popularity is recalculated, and the data in the cache is replaced according to the popularity ranking to ensure that the most likely to be accessed again hot files are always stored in the cache. The hot files will be stored in the Redis cache module. The Redis cache module is a Key-Value type database, where the Key value is the complete path of the stored file, and the Value value is the stored file content. When retrieving data, you can first access the Redis cache module through the rowkey to check if there is data caching. If it exists, directly access the Redis cache to return the corresponding data. If it does not exist, access the data in HDFS.
[0086] In a specific embodiment, the data merging method adopted by the data merging module is as follows:
[0087] First, analyze the historical access logs of small files to obtain the access situation of each small file by users;
[0088] Calculate the overall relevance of a small file to all files in a certain merge queue;
[0089] When the overall relevance reaches the preset merge threshold, merge the small file into the merge queue and store it in HDFS by calling the API of MapFile provided by Hadoop; the file structure of MapFile consists of two parts, index and data, which are used to store index files and data respectively.
[0090] Among them, the calculation of the overall relevance of a small file to all files in a certain merge queue includes:
[0091] Correlation(f, fi) = {(g, fi)|P(fi|g) ∈ (th1, th2) * P(ffi) ∈ (th3, th4)} (1)
[0092]
[0093] Among them, COrrelation(f, fi) in Formula 1 represents the degree of association between file f and file fi; (f, fi) represents the relational binary group composed of f and fi; P(fi|f) represents the probability that file fi is accessed within a period of time after file f is accessed; P(ffi) represents the probability that the user accesses both f and fi within a period of time; th represents different thresholds, and the value range is within the interval (0, 1).
[0094] In Formula 2, Correlation(f, q) represents the overall correlation degree between the current small file f and all files in a certain merge queue q; n represents the number of merge queues.
[0095] The data structures required by the merging method are described as follows:
[0096] Small file f: represents the current small file to be stored.
[0097] Merge queue q: responsible for temporarily storing small files. When the size reaches the merging threshold, it is merged into a MapFile.
[0098] Full set of merge queues qList: the full set of merge queues q, containing n qs.
[0099] Set of queues in use inuseList: a subset of qList, representing the set of qs that are non-empty and have not reached the merging condition.
[0100] fitList: a subset of inuseList, representing the set of qs in inuseList where the remaining space is greater than the current small file f.
[0101] candidateList: a subset of fitList, representing the set of qs in fitList where the relevance to the small file f exceeds the standard.
[0102] The specific steps of the merging algorithm are as follows, as Figure 3 shown:
[0103] a. Initialize the full set of merge queues qList with a capacity of n.
[0104] b. For the current small file f, first check inuseList. If inuseList is non-empty, select all eligible qs (i.e., qs with remaining space that can accommodate the current f) in it and add them to fitList; if inuseList is empty, f enters any empty q, and f enters inuseList.
[0105] c. If fitList is not empty, traverse the queues in fitList. For each queue qi, calculate the average correlation Correlation(f, q) of f in q. If Correlation(f, q) > fitThreshold, then f enters q; if Correlation(f, q) < fitThreshold, then q is removed from fitList and enters candidateList.
[0106] d. The setting of candidateList is to balance the two factors of the correlation between files and the file distribution. If only the correlation is considered, when the overall correlation of the files is not so close, it will cause waste of space. At this time, a compromise method is needed: select the queue q with the highest average correlation for file f in candidateList m , and judge Correlation(f, q m ). If the value exceeds SecondThreshold, then let f enter q m ; otherwise, the correlation is no longer considered, and f enters a new queue.
[0107] e. If fitThreshold is empty, it means that there is no merge queue that can accommodate the current small file. Check the queue q with the smallest remaining space in inuseList m , if the size of q m reaches the merge threshold (128MB of the HDFS block size), then merge the files in q into a MapFile.
[0108] For the merge queues that meet the merge conditions, the small files in them will be merged and stored in the form of MapFile provided by Hadoop. The file structure of MapFile consists of two parts: index and data, which are used to store the index file and data respectively. The key value of each record is the file name and is stored as an index by the index; the value stores the file content. Compared with the other two merge storage methods provided by Hadoop, MapFile has higher access efficiency.
[0109] In a specific embodiment, the method further includes a file writing process, and the file writing process includes:
[0110] Receive a write file request sent by the user, and determine whether the written file is a small file;
[0111] If it is a small file, enter the data merge module for data merge, and transfer the merged small file to HDFS for storage;
[0112] If it is not a small file, it directly enters the HDFS system for storage operations. Specifically, a write request is sent to the NameNode in the HDFS system. After receiving the file write signal, the NameNode searches for the corresponding DataNode and stores the file in the DataNode. The DataNode receives and confirms the data information for writing the file initiated by the NameNode, and finally returns a write report to the NameNode. The write result is returned to the user, and then the data stream is closed and the write operation is completed. The specific process is shown in Figure 4 shown below.
[0113] In a specific embodiment, the method further includes a file deletion process, and the file deletion process includes:
[0114] Receiving a file deletion request sent by the user; determining whether the file to be deleted is a small file;
[0115] If not, it directly enters the HDFS system for deletion operations;
[0116] If so, query the Redis cache module according to the complete file path; if there is a file to be deleted in the Redis cache module, perform a deletion operation on the small file cached in the Redis cache module;
[0117] If there is no file to be deleted in the Redis cache module, send a deletion request to the NameNode in the Hadoop cluster; the NameNode matches and interacts with the DataNode to find the corresponding small file, sends a deletion request for deletion operations; and returns the information of deleting the small file to the user. The specific process is shown in Figure 5 shown below.
[0118] Preferably, the method further includes:
[0119] When an update, deletion, or addition operation occurs to a small file in Hadoop, obtain the corresponding information through a message queue or an event listening mechanism, and capture these operation events in a timely manner;
[0120] The Redis cache module performs synchronous update, deletion, or addition operations on the small files cached in the Redis cache module according to the received information; thereby ensuring the consistency of the cached data and the data in the Hadoop storage system, and avoiding read errors caused by data inconsistency.
[0121] In this embodiment, the data in the cache can be comprehensively checked for integrity and repaired regularly (e.g., weekly). By comparing with the original data in the Hadoop storage system, it is detected whether there are any situations of data loss, corruption, or inconsistency in the cached data. If problems are found, corresponding repair measures are taken in a timely manner, such as reloading data from the Hadoop storage system, updating the incorrect data in the cache, etc., to ensure the accuracy and availability of the cached data.
[0122] This embodiment can also deeply analyze and tune the cache management strategy according to the system operation logs and performance monitoring data. For example, by analyzing metrics such as cache hit rate and disk I / O load in different time periods, the prefetch time window is optimized, and parameters such as the heat calculation period and decay coefficient are adjusted to improve the utilization rate of the cache space and the overall performance of the system, so that the system can always maintain efficient and stable operation in the scenario of reading a large number of small files.
[0123] Through the synergistic effect of multi - cache optimization, double elimination mechanism, and cache pre - loading strategy, this embodiment significantly improves the efficiency of reading a large number of small files based on Hadoop. In actual tests, compared with the traditional Hadoop reading method, the reading latency is reduced by more than 80%, and the system response speed is increased by 3 to 10 times. In the scenario of real - time analysis of financial transaction data, it can quickly read a large number of transaction detail small files, providing timely and accurate data support for risk assessment, market trend analysis, etc., and meeting the strict requirements of the business for real - time performance.
[0124] The present invention effectively improves the utilization rate of the cache space by combining the cache update strategy with the small file access prediction module, reducing the redundancy and invalid occupation of the cached data. By dynamically eliminating files with lower heat values according to file heat and timeliness, it ensures that the most valuable hot data is always stored in the cache, reducing the storage cost and resource consumption of the system. At the same time, it reduces the disk I / O operations in HDFS, reduces the disk load, extends the service life of the disk, further improves the overall performance and stability of the system, and enhances the competitiveness of the Hadoop system in the scenario of processing a large number of small files.
[0125] Obviously, the above - mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for reading massive small files in Hadoop based on file merging and dual elimination mechanisms of popularity and timeliness, the method being applicable to an HDFS system having a data merging module and a Redis cache module, and characterized by: The method comprises the following steps: Receive a small file read request input by a user, and determine whether the file requested for reading is a small file; If not, a request is sent to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information; If so, query the Redis cache module according to the complete path of the small file. If the Redis cache module caches the small file requested to be read, directly return the read small file to the user; the Redis cache module caches some hot small files according to the cache update strategy. The hot small files are predicted by the small file access prediction module constructed by the heat calculation formula based on the double elimination mechanism of hot spots and timeliness; If there is no small file requested to be read in the Redis cache module, a request is sent to the NameNode in the HDFS system, and the NameNode reads the corresponding file from the DataNode according to the metadata information.
2. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1 is characterized by: The heat calculation formula of the dual elimination mechanism based on hot spots and timeliness includes the following: Where α represents the attenuation coefficient, 0<α<1; countPeriod represents the heat calculation period; visitCount represents the number of times the small file is visited in the current heat calculation period; heat n-1 Indicates the historical popularity of small files.
3. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1 is characterized by: Before receiving a small file read request input by a user, the method further includes: initializing the HDFS system; The initialization includes: Configure network parameters and cache parameters, wherein the cache parameters include cache capacity and cache expiration time; Start the cache update strategy engine and load the pre-trained small file access prediction model; initialize relevant parameters and variables, including the historical data storage location and the small file access prediction model update frequency; Initialize the Redis cache module, set the initial size of the cache space, and configure relevant parameters of the cache update strategy; the relevant parameters of the cache update strategy include the preloading time window and the preloading file type.
4. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1 is characterized by: The cache update strategy includes: When executing a read request, only the number of times the file is accessed within the heat calculation period is recorded until the number of requests reaches countPeriod, that is, when the heat calculation period ends, the service process triggers the cache update and replacement operation; At the same time, the heat values of all files are calculated according to the heat calculation formula, and they are sorted according to the heat values. The top-K heat value sorted files and their metadata are cached in the Redis cache module using the "insert upon access" strategy; When the cache space of the Redis cache module is full, the file with the smallest heat value will be selected and eliminated from the cache according to the heat value of the file, and the files ranked TOP-K in heat value will be retained in the cache, thereby realizing dynamic update and optimization of cache data.
5. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 4 is characterized by: The cache update strategy also includes: utilizing the system idle period or low-load period to pre-load the files ranked TOP-K by heat value into the Redis cache module.
6. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1, characterized in that: The data merging method adopted by the data merging module is as follows: First, analyze the historical access logs of small files to obtain the user's access to each small file; Calculate the overall relevance of a small file to all files in a merge queue; When the overall relevance reaches a preset merge threshold, the small files are merged into a merge queue and stored in HDFS by calling the MapFile API provided by Hadoop; The file structure of the MapFile consists of two parts: index and data, which are used to store index files and data respectively.
7. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 6, characterized in that: The calculation of the overall relevance of a small file to all files in a merge queue includes: Correlation(f,fi)={(f,fi)|P(fi|f)∈(th1,th2)*P(ffi)∈(th3,th4)}(1) In formula 1, Correlation(f, fi) represents the correlation between file f and file fi; (f, fi) represents the relationship tuple consisting of two files f and fi; P(fi|f) represents the probability that file fi is accessed within a period of time after file f is accessed; P(ffi) represents the probability that a user accesses both f and fi within a period of time; th represents different thresholds, and the value range is within the interval (0,1); In Formula 2, Correlation(f,q) represents the overall correlation between the current small file f and all files in a certain merge queue q; n represents the number of merge queues.
8. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 6, characterized in that: The method further includes a file writing process, and the file writing process includes: Receive a file writing request sent by a user and determine whether the file being written is a small file; If it is a small file, it will enter the data merging module to merge the data and transfer the merged small file to HDFS for storage; If it is not a small file, it will directly enter the HDFS system for storage operations.
9. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1, characterized in that: The method further includes a file deletion process, wherein the file deletion process includes: Receive a file deletion request sent by a user; determine whether the deleted file is a small file; If not, directly enter the HDFS system to perform the deletion operation; If so, query the Redis cache module according to the complete file path; if there is a file requested to be deleted in the Redis cache module, delete the small files cached in the Redis cache module; If there is no file requested to be deleted in the Redis cache module, a deletion request is sent to the NameNode in the Hadoop cluster; the NameNode matches and interacts with the DataNode to find the corresponding small file, and sends a deletion request to perform the deletion operation; and the deletion information of the small file is returned to the user.
10. The method for reading massive small files in Hadoop based on file merging and dual elimination mechanism of popularity and timeliness according to claim 1, characterized in that: The method further comprises: When small files in Hadoop are updated, deleted, or added, the corresponding information is obtained through the message queue or event monitoring mechanism; The Redis cache module synchronously updates, deletes or adds small files cached in the Redis cache module according to the received information.
Citation Information
Patent Citations
Mass small file storage method based on Redis and HDFS
CN112650711A
Cited By
Method for eliminating and maintaining hotspot data in distributed cache system
CN120470036A