Column storage format and hierarchical index optimized Hadoop mass small file reading method

By adopting columnar storage format and hierarchical index optimization methods in the Hadoop platform, combining the hotspot cache index layer and the persistent index storage layer, the problem of inefficiency in reading massive small files in traditional Hadoop is solved, and fast and accurate data query and system stability are achieved.

CN120353759APending Publication Date: 2025-07-22GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316534.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional Hadoop method is inefficient when reading massive small files, resulting in a sharp increase in NameNode memory pressure and low reading efficiency, affecting system performance and stability.

Method used

The columnar storage format and hierarchical index optimization method is adopted. Through the combination of the hotspot cache index layer and the persistent index storage layer, the hotspot data is first queried in memory. If there is no persistent index storage layer, the structure of the HBase original table and index table is optimized, and multi-attribute comprehensive query is supported.

Benefits of technology

It significantly improves the reading efficiency of massive small files, shortens query response time by 5 to 10 times, reduces system resource consumption, enhances system performance and stability, and adapts to the growing data demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353759A_ABST
    Figure CN120353759A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computers, and discloses a column storage format and hierarchical index optimized Hadoop mass small file reading method, which is suitable for building an Hbase system for storing mass picture small files and attribute information thereof in a Hadoop platform, and a distributed coordination management system ZooKeeper. The method comprises the following steps: initiating a query request to a service process of a hotspot cache index layer based on a memory, and querying whether hotspot data cached in the hotspot cache index layer has query request data or not according to a consistent Hash algorithm; if yes, directly returning a query result fed back by the hotspot cache index layer; if the persistent index storage layer does not exist, the query is forwarded to the persistent index storage layer based on the Hbase for query, and a query result is returned; the persistent index storage layer stores an HBase original table which is established based on an HBase picture storage strategy and is used for storing small picture files and attribute information of the small picture files, and an index table which is established according to non-row key attributes in the HBase original table.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method for reading a large number of small files in Hadoop with a columnar storage format and hierarchical index optimization. Background Art

[0002] With the rapid development of Internet, cloud computing, and big data technologies, the amount of data generated by various industries shows an exponential growth trend. In the field of intelligent monitoring, the cameras densely distributed in cities generate tens of thousands or even millions of small picture files every day; on e-commerce platforms, product display pictures, pictures uploaded by users, etc. also constitute a large amount of small file datasets; in geographic information systems, satellite images, aerial pictures, etc. are also stored in large quantities in the form of small files.

[0003] Therefore, the demand for the storage and processing of massive data in many key fields such as intelligent monitoring, e-commerce, and geographic information systems has increased explosively. Among them, a large number of small picture files, as an important part of the data, the storage and reading efficiency thereof is directly related to the performance of these applications and the user experience.

[0004] These massive small files face many severe challenges under the traditional Hadoop Distributed File System (HDFS) storage architecture, which are roughly as follows:

[0005] 1. Dramatic increase in NameNode memory pressure: In HDFS, the NameNode is responsible for managing the metadata of all files, including information such as the name, size, and storage location of the files. Each small file needs to occupy a certain amount of memory space in the NameNode to store this metadata. When the number of small files is huge, the memory consumption of the NameNode will increase sharply. For example, if the metadata of each small file occupies an average of 1 KB of memory, 1 million small files will occupy nearly 1 GB of memory. This will not only lead to a decline in the performance of the NameNode, but may even cause a memory overflow error, making the entire HDFS system unable to work properly.

[0006] 2. Low reading efficiency: The reading efficiency of small files is severely restricted in HDFS. Due to the small data volume of small files, multiple disk addressing operations are required for each read. Disk addressing is a relatively time-consuming process, and frequent disk addressing operations will greatly increase the reading latency. Taking the example of reading 100 small files of 100 KB each, compared with reading a large file of 10 MB, the number of disk addressing operations may increase by dozens of times, resulting in a significant extension of the reading time and seriously affecting the response speed and data processing ability of the system. Summary of the Invention

[0007] To overcome the problem of low efficiency in reading a large number of small files by the traditional Hadoop method, the present invention provides a method for reading a large number of small files in Hadoop with a columnar storage format and hierarchical index optimization.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0009] A method for reading a large number of small files in Hadoop with a columnar storage format and hierarchical index optimization, the method is applicable to a Hbase system for storing a large number of small picture files and their attribute information, and a distributed coordination management system ZooKeeper built on the Hadoop platform. The reading method includes the following:

[0010] Initiate a query request to the service process of the in-memory hot cache index layer, and query whether the hot data cached in the hot cache index layer exists for the query request data according to the consistent hashing algorithm;

[0011] If it exists, directly return the query result feedback by the hot cache index layer;

[0012] If it does not exist, forward the query to the persistent index storage layer based on Hbase for query, and return the query result; the persistent index storage layer stores an HBase original table established based on the picture storage strategy of HBase for storing small picture files and their attribute information, and an index table established according to the non-row key attributes in the HBase original table.

[0013] Preferably, before initiating a query request to the service process of the in-memory hot cache index layer, the method further includes: obtaining the address of ZooKeeper from the configuration file, establishing a connection with ZooKeeper, and obtaining all registered service processes; determining the location information of all service processes currently providing memory caching.

[0014] Preferably, establishing an HBase original table based on the picture storage strategy of HBase for storing small picture files and their attribute information includes:

[0015] Establish an HBase original table, where one column of the column family in the HBase original table stores small picture files, and the small picture files are stored in this column in the form of binary data;

[0016] Other columns of the column family in the HBase original table store various attribute information of the pictures, and the attribute information includes small file type, size, creation time, and modification time.

[0017] Further, the HBase original table adopts a column-oriented storage model to continuously store data of the same column family. When storing each column family, the data in each row cell is stored in the form of Key-Value to form several data blocks. Then, the data blocks are saved into an HFile, and finally the HFile is saved onto the background HDFS.

[0018] Preferably, after the index table established according to the non-row key attributes in the HBase original table, the method further includes: saving the index table in the HBase system to implement persistent storage of the index data for the HBase original table; each index table is used to store the index of a certain non-row key attribute to be queried in the HBase original table.

[0019] The format of the primary key of the index table defined for the non-row key attribute to be indexed in the HBase original table is as follows: <original table index attribute name, original table index attribute value, original table row key>.

[0020] Preferably, when the query request includes multiple non-row key attributes at the same time, the index table established according to the non-row key attributes in the HBase original table includes:

[0021] Construct a combined index table for multiple non-row key attribute columns;

[0022] Through the established combined index table, convert the combined query into a query based on the primary key of the index table;

[0023] Save the combined index table in the HBase system to implement persistent storage of the index data for the HBase original table; each combined index table is used to store the index of a certain non-row key attribute to be queried in the HBase original table.

[0024] The format of the primary key of the combined index table defined for the non-row key attribute to be indexed in the original table is as follows: <original table index attribute name, original table index attribute value, original table row key>.

[0025] Preferably, the hot data cached in the hot cache index layer is obtained through the following steps:

[0026] Based on the law that the access to network data satisfies the Pareto distribution, screen the user access log records on the HBase system, sort them according to the access volume, and select the top several users with the highest access volume as active users;

[0027] Use the log-linear algorithm as the index hot data prediction model to calculate the hot value of the access data of each active user, and mark the picture small files with the top K% of the hot values as hot data; where K represents the hot threshold calculated according to the limit of the number of records that can be accommodated in the cache space.

[0028] Cache the hot data in the hot cache index layer in a preset format.

[0029] Furthermore, the calculation formula of the index hot data prediction model is as follows:

[0030] ln N i = k(t)ln N i (t) + b(t)

[0031] Wherein, N t is the predicted total access volume of file i, that is, the hot value; N i (t) represents the access volume of file i within the observation time, and the length of the observation time is t; k(t) and b(t) are the relevant parameters of the linear relationship, and the optimal values are calculated using the linear regression method; the length of the observation time t represents the time difference between the access start time element of the record line in the user access log record and the time when the user access log record is collected.

[0032] Furthermore, the format of the hot data cached in the hot cache index layer is as follows:

[0033] Index primary key: <original table index column name, original table index column value>

[0034] Index set: {<original table primary key, {<frequently accessed column name, frequently accessed column value>}>}.

[0035] Preferably, the reading method further includes: when a new picture is written or an existing picture is modified or deleted, an incremental update method is adopted to only update the index related to the changed picture.

[0036] Compared with the prior art, the beneficial effects of the present invention are:

[0037] The reading method of the present invention first queries whether the query request data exists in the hot data cached in the hot cache index layer. If it exists, the query result feedback by the hot cache index layer is directly returned. This reduces disk I / O operations and data processing volume, and reduces the resource consumption of the system. At the same time, the present invention first queries the hot cache index layer. If not, it then queries based on the persistent index storage layer of Hbase. This hierarchical index optimization makes the search for small picture files faster and more accurate, effectively reducing the reading latency. At the same time, the method of the present invention has good scalability and fault tolerance, can adapt to the growing storage and query requirements of a large number of small files, enhances the overall performance and stability of the system, and provides a reliable technical guarantee for big data applications in various industries.

[0038] The present invention significantly improves the reading efficiency of massive small files based on Hadoop. In actual tests, compared with the traditional Hadoop small file reading method, for hot queries, due to cache hits, the query response time is shortened by 5 to 10 times, greatly reducing the time for reading the same number and size of small image files. This enables quick response to user query requests in application scenarios such as intelligent monitoring and e-commerce, improving the real-time performance of the system and the user experience. Brief Description of the Drawings

[0039] Figure 1 is a flowchart of the steps of the method for reading massive small files of the columnar storage format and hierarchical index optimization of Hadoop according to the present invention.

[0040] Figure 2 is a schematic diagram of the hierarchical index structure of the present invention.

[0041] Figure 3 is an example diagram of the index structure of the hot cache index layer of the present invention.

[0042] Figure 4 is a schematic diagram of the storage node mapping of the consistent hashing algorithm of the present invention.

[0043] Figure 5 is a schematic diagram of the storage node migration of the consistent hashing algorithm of the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. The present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0045] It should be understood that when used in this specification, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0046] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0047] It should be further understood that the term "and / or" used in the specification of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0048] Example 1

[0049] To overcome the problem of low efficiency in reading a large number of small files by the traditional Hadoop method mentioned in the background art, the present invention proposes a method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization. It innovatively combines an advanced columnar storage format with a unique hierarchical index structure to comprehensively optimize the method for reading a large number of small files based on Hadoop. Through this deeply integrated method, it aims to fundamentally improve the reading efficiency, significantly reduce system resource consumption, enhance the performance and stability of the system in the scenario of processing a large number of small files, and meet the urgent needs of various industries for efficient big data processing.

[0050] As Figure 1 described, a method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization provided by the present invention is applicable to a Hbase system for storing a large number of small picture files and their attribute information and a distributed coordination management system ZooKeeper built on the Hadoop platform. The reading method includes the following:

[0051] Initiate a query request to the service process of the in-memory hot cache index layer, and query whether the hot data cached in the hot cache index layer exists for the query request data according to the consistent hashing algorithm;

[0052] If it exists, directly return the query result feedback by the hot cache index layer;

[0053] If it does not exist, forward the query to the persistent index storage layer based on Hbase for query, and return the query result; the persistent index storage layer stores an HBase original table established for storing small picture files and their attribute information based on the picture storage strategy of HBase, and an index table established according to the non-row key attributes in the HBase original table.

[0054] The present invention first queries the hot index data in the hot cache index layer according to the query request. If it is not hit in the cache (that is, there is no data related to the query request), the query is forwarded to the persistent index storage layer for retrieval. It can be seen that by caching the hot data in memory, some queries can be directly hit in memory, thereby reducing the disk access overhead and improving the overall query performance, which is particularly effective for applications with skewed data access distribution characteristics.

[0055] In a specific embodiment, before initiating a query request to the service process of the in-memory hot cache index layer, the method further includes: obtaining the address of ZooKeeper from a configuration file, establishing a connection with ZooKeeper, and obtaining all registered service processes; determining the location information of all service processes currently providing memory caching.

[0056] In a specific embodiment, since the present invention selects the HBase system as the platform for storing a large number of small image files and their attribute information. The HBase system is a simple structured data distributed storage technology based on HDFS, which can be used to store a large number of small image files and has various advantages such as system-level small file merging and global namespace. At the same time, as a column-oriented distributed unstructured database, the HBase system can provide fast random access to a large amount of unstructured data (images, texts, audios, etc.), and has characteristics such as high reliability, high performance, column storage, and scalability. Its column storage method is essentially different from the traditional row storage method. In traditional row storage, data is stored in units of rows, and each row of data contains the values of multiple fields. In the column storage of HBase, however, data is organized by column families, and the data in the same column family is physically stored together. This enables, when querying, if only certain columns or a few columns of data need to be obtained, there is no need to read the entire row of data, greatly reducing the amount of data read and improving the query efficiency.

[0057] Establish an HBase original table for storing small image files and their attribute information based on the HBase-based image storage strategy, including:

[0058] Establish an HBase original table, where one column of the column family in the HBase original table stores small image files, and the small image files are stored in this column in the form of binary data;

[0059] The other columns of the column family in the HBase original table store various attribute information of the images, and the attribute information includes small file type, size, creation time, and modification time; the attribute information may also include information related to specific applications.

[0060] In this embodiment, the HBase original table adopts a column-oriented storage model to continuously store data of the same column family; when storing each column family, the data in each row of cells is stored in the form of Key-Value to form several data blocks; then the data blocks are saved to an HFile, and finally the HFile is saved to the underlying HDFS.

[0061] In this embodiment, since the content of the small picture file is stored in a cell, the process of packing the small picture file is actually implied during the data storage process. However, due to the data block limit in HBase, it is necessary to adjust according to the application. By default, the HBase data block limit is 64 KB. Since the picture content is saved as the value of a cell, its size is restricted by the size of the data block. In the application, the HBase data block size needs to be modified according to the maximum picture size. The specific modification method is to specify the data block size with HColumnDescriptor when creating the original table, which can be specified for each column family. By storing the picture attribute information and the picture content in a large original table, the comprehensive query of multiple attributes of the picture can be supported. In addition, according to the application requirements, the column family can be extended to save application-related information, so as to support the picture query related to the application. For example, in the urban traffic monitoring system, in addition to the above standard attributes, the vehicle license plate information identified from the picture, the ID of the shooting camera, etc. can be stored as application-related attributes in the corresponding columns. By storing the picture attribute information and the picture content in a large table, the comprehensive query of multiple attributes of the picture is realized. Taking the query of a picture with a certain license plate number as an example, the system only needs to read the data in the column storing the "license plate number information" to quickly locate and obtain the corresponding picture, without traversing the massive pictures in the entire HDFS, thus significantly improving the reading performance and data processing efficiency. This storage strategy not only effectively solves the problem of excessive memory pressure on the NameNode, but also provides strong support for flexible picture retrieval. In addition, HBase implies the small file packing process and realizes the small file merging without secondary development. HBase uses a distributed B+ tree to globally and uniformly manage the picture metadata, realizing a global namespace and facilitating the management of pictures.

[0062] In a specific embodiment, although storing a large number of small image files into the HBase raw table not only solves the memory pressure of the NameNode in HDFS, but also supports the comprehensive query of multiple attributes of images. However, the HBase raw table only has a primary key index and does not support non-primary key indexes, which results in a low data query efficiency of the HBase raw table and is difficult to meet the requirements of real-time or near-real-time data query. Retrieving data in the HBase raw table usually has three methods: specifying a single row key query, specifying a range query of row keys, and a scan operation. Among them, the scan operation is mainly used for querying non-row key attributes. Although it allows specifying the data range to be scanned and specifying conditional constraints during the query process to filter the target result set, which has high flexibility, the scan operation for non-row key attributes is extremely costly, and its time complexity is O(N). While the time complexity of retrieving based on row keys is only O(logN). In order to improve the efficiency of full table scanning during non-primary key query in HBase, the present invention establishes an index table for non-row key attributes stored in the HBase raw table.

[0063] Specifically, after the index table established according to the non-row key attributes in the HBase raw table in this embodiment, the method further includes: storing the index table in the HBase system, and leveraging the good scalability and fault tolerance of HBase to improve the performance of the index, so as to achieve the persistent storage of index data for the HBase raw table; each index table is used to store the index of a certain non-row key attribute to be queried in the HBase raw table;

[0064] The format of defining the primary key of the index table for the non-row key attribute to be indexed in the HBase raw table is as follows: <raw table index attribute name, raw table index attribute value, raw table row key>.

[0065] In another specific embodiment, in the actual application scenario, there is often a need for combined query of multiple non-row key attributes. Similar to the multi-field index in a database, the present invention constructs a combined index table of multiple non-row key attribute columns to meet this need.

[0066] When the query request includes multiple non-row key attributes at the same time, the index table established according to the non-row key attributes in the HBase raw table includes:

[0067] Constructing a combined index table of multiple non-row key attribute columns;

[0068] Through the established combined index table, converting the combined query into a query based on the primary key of the index table;

[0069] Storing the combined index table in the HBase system to achieve the persistent storage of index data for the HBase raw table; each combined index table is used to store the index of a certain non-row key attribute to be queried in the HBase raw table;

[0070] Define the format of the primary key of the composite index table for the non-row key attributes to be indexed in the original table as follows: <Original table index attribute name, Original table index attribute value, Original table row key>.

[0071] For example, when the conditions of the query request contain multiple non-row key attributes at the same time, such as in the urban traffic monitoring system, when querying "vehicle number" and "shooting camera ID" at the same time, an index table with these two non-row key columns as the primary key is established. The primary key of the index table is in the form of "license plate number, Yue E2051D, shooting camera ID, 1, image1". After establishing a composite index table on multiple attribute columns, the composite query can be converted into a query based on the primary key of the index table, and its indexing process and query process are basically the same as those of a single-attribute query, thus greatly improving the efficiency of the composite query.

[0072] The reading method described in the present invention is divided into the following functional layers according to functions, and these functional layers constitute the entire hierarchical index storage system, as Figure 2 shown:

[0073] (1) Index construction management layer. Manage the metadata of the index (record information such as the index table name corresponding to the original table, index columns, etc.), and implement index construction methods for two different types of data, namely streaming data and static data, for HBase, including supporting insert, delete, and update operations on the index table and value table.

[0074] (2) Persistent index storage layer. Provide persistent storage for the index table, and HBas provides scalability and fault tolerance for persistent storage of data.

[0075] (3) Hot cache index layer. Manage the cache storage, update, and address mapping of index hot data, so that the data frequently accessed recently can be cached in memory.

[0076] (4) Query execution engine. Translate the user's query request into a command recognized by the system, call the corresponding method to execute the query, and summarize and return the query result to the client.

[0077] To improve the high availability of the index memory cache, the distributed coordination management system ZooKeeper in the Hadoop environment is used to detect the survival status of service processes on distributed memory nodes. Each memory node service process in the hot cache index layer will establish a session with ZooKeeper respectively and create a temporary Znode to represent its own survival status. Each memory node service process can observe the survival status of other node processes from the ZooKeeper system image. Through the monitoring of the status of distributed memory nodes, the failure detection and online processing of memory nodes are realized, so as to realize the high availability of the hot cache index layer of the index.

[0078] Table 1 and Table 2 show the relationship between the HBase original table and the persistent index table of HBase. Each row record in the HBase original table has a unique row key and multiple attribute columns. The key part of the primary key of the index table consists of the non-row key attributes to be queried in the original table, the attribute values, and the row key of the original table. Storing the row key of the original table in the primary key of the index table has two functions: one is to ensure the uniqueness of the primary key of the index table; the other is to provide the address of the indexed record in the HBase original table. Through the row key of the original table, the indexed record in the original table can be quickly obtained. The Value part stores the attribute values that need to be accessed in the original table. Through this index structure, the data that meets specific conditions in the original table can be quickly located and obtained, improving the query efficiency.

[0079] Table 1 HBase Original Table

[0080] Row key Attribute A Attribute B Attribute C … Row key value Attribute A value Attribute B value Attribute C value …

[0081] Index Table Corresponding to the HBase Original Table:

[0082] Table 2 Index Table

[0083]

[0084] In a specific embodiment, the hot data is based on the following assumption: If the access volume of users is observed over several time periods (such as 24 hours) to predict the total access volume of the data by users in the next step, then obviously there is the following rule: The more times the data is accessed by users during the observation time, the more likely the data can be regarded as hot data, and the greater the possibility that the data will be accessed by users next time.

[0085] Therefore, the hot data cached in the hot cache index layer is obtained through the following steps:

[0086] Based on the rule that the access of network data satisfies the Pareto distribution, the user access log records on the HBase system are screened, sorted according to the access volume, and the top users with the most access volumes are selected as active users. The rule that the access of network data satisfies the Pareto distribution means that most I / O requests access a small amount of hot data, and the access volume of 20% of users accounts for about 80% of the total access volume of all users, and most of the 80% of the access volume is concentrated on 20% of the data. Therefore, the user access log records on the HBase original table are screened, and the top 20% of users with the most access volumes are selected as active users.

[0087] Use the log-linear algorithm as the index hot data prediction model to calculate the hot value of each active user's access data, and mark the picture small files with the top K% of the hot values as hot data; where K represents the hot threshold calculated according to the limit of the number of records that can be accommodated in the cache space.

[0088] Cache the hot data in the hot cache index layer in a preset format.

[0089] In this embodiment, the calculation formula of the index hot data prediction model is as follows:

[0090] ln N i = k(t)ln N i (t) + b(t)

[0091] In the formula, N t is the predicted total access volume of file i, that is, the hot value; N i (t) represents the access volume of file i within the observation time, and the length of the observation time is t; k(t) and b(t) are the relevant parameters of the linear relationship, and the optimal values are calculated using the linear regression method; the length of the observation time t represents the time difference between the access start time element of the record row in the user access log record and the time when the user access log record is collected.

[0092] In a specific embodiment, the index table will implement the persistent storage of index data for the Hbase original table. Since the index data is stored in the Hbase original table, each query access to the Hbase original table will involve a lot of disk accesses. Further, cache the index data with high access frequencies in the index as hot data in the memory to form a hierarchical index storage and query mechanism based on Hbase and distributed memory, further improving the query speed of the index. The index format of the hot data cached in the hot cache index layer is different from the index format in the persistent storage. The primary key format of the in-memory cache index in the hot cache index layer is:

[0093] 〈Original table index column name, original table index column value〉.

[0094] Among them, the meanings of the original table index column name and the original table index column value are the same as those in the persistent index storage layer. Each index primary key in the hot cache index layer corresponds to a set of index records with the same index column value, and this set contains all the index table data records corresponding to this index value. Similar to the persistent index storage layer, the set also contains other non-row key attributes that may need to be accessed. Therefore, the format of the hot data cached in the hot cache index layer is as follows:

[0095] Index primary key: 〈Original table index column name, original table index column value〉

[0096] Index set: {〈Original table primary key, {〈Frequently accessed column name, frequently accessed column value〉}〉}.

[0097] The present invention completes the storage management of index hot data in distributed memory by introducing the consistent hashing algorithm. The consistent hashing algorithm provides good scalability for the hot cache index layer of Hbase. The scalability of the hot cache index layer means that when the memory utilization rate of the hot cache index layer is relatively high, the capacity of the hot cache index layer can be dynamically increased by adding new service nodes. In a dynamically changing memory environment, monotonicity is a reliable guarantee for scalability. Monotonicity means that if some content has been assigned to the cache of the corresponding node through the hashing method, and then new nodes are added to the system, the hashing result should be able to ensure that the original assigned content is either mapped to the cache of the new node or still mapped to the cache of the original node, so as to minimize the overhead caused by data migration. For example, for the simplest linear hashing:

[0098] Address=ax + b mod(N)

[0099] Among them, N represents the number of all buffer areas, that is, the number of nodes in the hot cache index layer. If the consistent hashing algorithm is not used to manage the distributed index cache, when the number of cache nodes N changes due to the addition and withdrawal of service processes, all the original hashing results will change, which means that all mapping relationships need to be updated in the system. The addition and withdrawal of nodes will bring great computational and transmission overheads.

[0100] In a specific embodiment, to ensure the accuracy and timeliness of the index, the reading method further includes: when new pictures are written or existing pictures are modified or deleted, an incremental update method is adopted to only update the index related to the changed pictures, avoiding the reconstruction of the entire index. For example, when a new picture is added to the HBase original table, the system will add a new index record to the corresponding index table according to its attribute information and the index construction rules; when a certain attribute of an existing picture changes, the system will locate the corresponding index record of the picture in the index table and update the relevant attribute values. This incremental update method greatly reduces the overhead of index maintenance and ensures that the index can always accurately reflect the state of the original data.

[0101] Figure 3The relationship between the HBase original table and the persistent index table of HBase is shown. Each row record in the HBase original table has a unique row key and multiple attribute columns. The key part of the primary key of the index table consists of the non-row key attributes to be queried in the original table, the attribute values, and the row key of the original table. Storing the row key of the original table in the primary key of the index table has two functions: one is to ensure the uniqueness of the primary key of the index table; the other is to provide the address of the indexed record in the HBase original table. Through the row key of the original table, the indexed record in the original table can be quickly obtained. The Value part stores the attribute values that need to be accessed in the original table. Through this index structure, the data that meets specific conditions in the original table can be quickly located and obtained, improving the query efficiency.

[0102] Tables 3 and 4 give an example of the HBase original table storing picture information and the persistent index table of HBase with the license plate number attribute as the index primary key. In this example, the primary key of the index table is like "license plate number, Yue E2051D, image1", where the license plate number is the non-row key attribute of the original table and also the non-row key attribute to be queried; Yue E2051D is the specific value of the license plate number of the image1 picture in the original table data, that is, the index attribute value; image1 is the row key corresponding to this record in the original table. Similar to some indexes in relational databases, the query attribute of HBase is the primary key of the index table. The index table contains some fields of the original table. Usually, only the fields that may need to be accessed assistively in the query are stored in the non-primary key attributes of the index table. This is to facilitate fast access to the assistive fields that need to be accessed in the query and avoid secondary disk access caused by accessing the original table again. In this example, the value saved in the value of the index table is {picture content, AAAIffG29 / / +23EF..}, indicating the picture content attribute that may need to be accessed when querying the license plate number. The above describes the situation of a single non-row key attribute index. In actual applications, there is a need for combined queries of multiple non-row key attributes. Therefore, similar to the multi-field index in the database, a combined index table of multiple non-row key attribute columns needs to be constructed. For the combined query situation of multiple non-row key attribute columns, HBase will establish a combined index based on multiple query attribute columns. For example, when the conditions of the query request include both "vehicle number" and "shooting camera ID" at the same time, a combined index table with these two non-row key columns as the primary key is established. The primary key of the combined index table is formed like "license plate number, Yue E2051D, shooting camera ID, 1, image1". After establishing indexes on multiple attribute columns, the combined query is converted into a query based on the primary key of the index table. The indexing process and the query process are basically the same as those of the single-attribute query.

[0103] Table 3 HBase original table

[0104]

[0105]

[0106] The index table corresponding to Table 3 is shown in Table 4

[0107] Table 4 Index Table

[0108]

[0109]

[0110] Figure 3 It is an example of the index structure of the hot spot cache index layer. In this example, a value "Yue E2051D" under the license plate number attribute is cached, corresponding to the primary key "license plate number, Yue E2051D" of the hot spot cache index layer. When querying with the license plate number Yue E2051D, the storage address of the corresponding index in memory can be found according to the hash value of the primary key "license plate number, Yue E2051D" of the index hot spot cache index layer, and the index set {〈image1, {picture content, AAAIffG29 / / +23EF..}〉 corresponding to the license plate number value can be obtained. The corresponding user data record can be obtained from the original table according to the primary key of the original table corresponding to this set. In implementation, the data of the hot spot cache index layer is stored in the Redis in-memory database, and Redis automatically completes the above-mentioned hashing and fast query processes. The all-memory query of this layer is quite efficient.

[0111] In this embodiment, the consistent hashing algorithm is adopted on the distributed memory of the HBase server node to manage all index hot spot data. The basic principle of the consistent hashing algorithm is as Figure 4 、 5 shown: (1) Construct a circular hash space. Use the hash function to map the value to the circle (continuum) from 0 to 2 32 . (2) Map the data object to the hash space. Calculate the hash value key through the hash function, and the key value must be distributed on the circular hash space. (3) Map the cache to the hash space. Use the same hashing algorithm to map both the object and the cache to the same hash numerical space. (4) Map the data object to the cache. Starting from the mapping point of the data object on the circular hash space and moving clockwise, the first cache node found is the storage location of the data object.

[0112] The present invention reduces the data transmission overhead caused by the change (increase or decrease) of cache nodes through the consistent hashing algorithm. For example, when a cache node fails, only the data objects between this cache node and the previous cache node need to be migrated, rather than all the stored data objects. As Figure 4When a failure occurs in the storage node cache2, the data objects originally mapped to cache2 will continue to be mapped to the node cache3 in the clockwise direction. Consistent hashing ensures the balance of each storage node by hashing data to different storage nodes. When querying the index information of data, the system will find the index information of the data in two steps: (1) Perform consistent hashing on the data, as Figure 5 shown, to find the storage node where the data index information is located; (2) Use the hashing mechanism (Redis) to find the index data address within the node.

[0113] Embodiment 2

[0114] Based on the method for reading a large number of small files in Hadoop with the columnar storage format and hierarchical index optimization described in the above Embodiment 1, in actual application, it is as follows:

[0115] First, build Hbase and ZooKeeper in the Hadoop system. ZooKeeper is a high-performance, distributed coordination service, which plays a crucial role in the whole system. ZooKeeper works in cooperation with Hadoop and HBase, responsible for managing node information in the cluster, coordinating distributed operations, etc. During the building process, it is necessary to carefully configure various parameters of Hadoop, HBase, and ZooKeeper to ensure their stable and efficient cooperation. For example, reasonably configure the file system parameters of Hadoop, optimize the storage and read / write parameters of HBase, and set the number of cluster nodes and election strategy of ZooKeeper, etc., to lay a solid foundation for subsequent picture storage, index construction, and data query.

[0116] Store a large number of small picture files and their attribute information in HBase. First, establish the HBase original table. According to the picture storage strategy described in Embodiment 1, store the picture content in one column of a separate column family, and store the application-related attribute information in the columns of other column families. Taking the pictures taken by the cameras in the urban traffic system as an example, store the vehicle license plate information, shooting time, shooting camera ID, etc. obtained by identifying the pictures in the corresponding columns. During the storage process, using the distributed characteristics of the HBase original table, the picture data can be evenly distributed and stored on multiple nodes, improving the reliability and scalability of storage. At the same time, through the batch writing function of HBase, the efficiency of picture storage can be improved and the writing time can be reduced.

[0117] According to the different data input methods, index construction can be divided into index construction for streaming data and index construction for batch data.

[0118] For the streaming data index construction method, the callback function prePut of the RegionObserver interface provided by HBase is used, and it will be triggered and called before inserting a Put data. The prePut method first analyzes the Put operation initiated by the user according to the index information. If the data of the Put operation contains index columns, that is, contains the data to be indexed, the insertion of the index data is triggered.

[0119] Since static data is generally relatively large, in order to accelerate the construction speed of static data indexes, the present invention uses Hadoop's MapReduce to perform the construction of static data indexes in parallel. The MapReduce task first inputs <Row, Result>, where Row is the row key of the original table, and Result is the HBase record obtained through Row. Then, according to the index information, the corresponding index data is generated, and the index data is updated to the persistent index storage layer and the hot cache index layer respectively, and the value table is updated. The entire process can be completed without the Reduce stage. The Hbase original tables are independent of each other, and MapReduce parallel processing is fully utilized to accelerate the index construction process.

[0120] The screening process for active users in the present invention is as follows:

[0121] (1) Use regular expression matching to screen out the record rows that access small picture files.

[0122] (2) Write a log parsing class to separately parse the five components of the record row, and convert the parsed record row into a bean object. The attributes of the bean object respectively correspond to the five components of the record row. Then call the JDBC API to persist the bean object into the MySQL database.

[0123] (3) Use a two-dimensional array to store the user IP and picture file name of the bean object.

[0124] (4) Traverse the user IP elements in the two-dimensional array, design a counter to count the access volume of each user IP. Use a Hash Map collection, with the user IP as the Key value and the Value value as the access volume of the user.

[0125] (5) Sort the Hash Map collection generated in step (4) in descending order according to the Value value, screen out the top 20% of the users and mark them as active users, and use an Array List collection to store the user IPs of the active users.

[0126] In the present invention, the hot value calculation process of the hot data is as follows:

[0127] (1) Write a nested for loop. The outer loop traverses and retrieves the user IP elements stored in the ArrayList collection (i.e., the ArrayList collection generated in step (5) of filtering active users), and the inner loop traverses and retrieves the user IP elements of the two-dimensional array (i.e., the array generated in step (3) of filtering active users), and the two are matched.

[0128] (2) After successful matching, use the user IP as the keyword and use the MySQL query statement to query the access start time of each user. Then calculate the hotness value of the data for each picture file accessed by the active users. Finally, sort in descending order according to the hotness value of the data, and mark the top K% of the files with the file hotness value as hot data. Finally, use the Redis database to cache the hot data. Redis organizes data in the <key, value> format. The index primary key of the hot data is used as the key, and the index set is saved as the value in the in-memory cache.

[0129] The query process described in the present invention is divided into single-value query and range query.

[0130] Among them, for single-value query, by establishing an index on non-primary key attributes, it supports efficient single-value query on non-row key columns.

[0131] The process of the client for single-value query is as follows:

[0132] 1. Obtain the address of ZooKeeper from the configuration file, establish a connection with ZooKeeper, obtain all registered service processes, and determine the location information of all service processes in the in-memory cache of the current hot cache index layer. This step is to be able to quickly locate the cache location storing the query data and improve the query efficiency.

[0133] 2. Send a query request to the service process of the hot cache index layer. If the hot cache index layer is hit, directly return the result given by the in-memory cache layer and end the query. The hot cache index layer usually stores hot data that is frequently accessed recently and can quickly respond to query requests, reducing query latency.

[0134] 3. If the hot cache index layer is not hit, send a query request to the index table of Hbase. After obtaining the result, return the query result and end the query. Through this way of first querying the cache and then querying the index table, the present invention makes full use of the high-speed characteristics of the cache and at the same time ensures that data can be accurately obtained when the cache is not hit.

[0135] The so-called range query is a query in which the value of the conditional attribute in the query request is a range. The process of the client for range query is as follows:

[0136] 1. Obtain the address of ZooKeeper from the configuration file, establish a connection with ZooKeeper, obtain all registered service processes, and determine the location information of all service processes that currently provide memory caching.

[0137] 2. According to the conditions of the range query, the client obtains all index column values existing between the ranges from the HBase index table.

[0138] 3. For all existing index column values, calculate the storage node address according to the consistent hashing algorithm, so as to correspond all existing index column values with the relevant node addresses one by one.

[0139] 4. Concurrently initiate query requests to the relevant nodes. Among them, multiple query requests initiated to the same node will be merged into a batch request.

[0140] 5. The in-memory index service processes on each node respond to the query requests. If the queried content is in memory, directly return the data in memory; otherwise, the service process will initiate a query to the persistent index storage layer and return the query result.

[0141] 6. The client aggregates the query results returned from each service node.

[0142] Through the adoption of the columnar storage format and hierarchical index optimization, the present invention significantly improves the reading efficiency of massive small files based on Hadoop. In actual tests, compared with the traditional Hadoop small file reading method, due to cache hits, the query response time of hot queries is shortened by 5 to 10 times, greatly shortening the time for reading the same number and size of picture small files. This enables quick response to users' query requests in application scenarios such as intelligent monitoring and e-commerce, improving the real-time performance of the system and the user experience.

[0143] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for reading a large number of small files in Hadoop with a columnar storage format and hierarchical index optimization, which is applicable to a Hadoop platform with a Hbase system for storing a large number of small picture files and their attribute information, and a distributed coordination and management system ZooKeeper, and is characterized in that: The reading method includes the following: Initiate a query request to the service process of the memory-based hot cache index layer, and query whether the hot data cached in the hot cache index layer exists for the query request data according to the consistent hashing algorithm; If it exists, directly return the query result feedback by the hot cache index layer; If it does not exist, forward the query to the persistent index storage layer based on Hbase for query, and return the query result; The persistent index storage layer stores an HBase original table established based on the HBase-based picture storage strategy for storing picture small files and their attribute information, and an index table established according to the non-row key attributes in the HBase original table.

2. The method for reading a large number of small files of Hadoop with columnar storage format and hierarchical index optimization according to claim 1, characterized in that: Before initiating a query request to the service process of the memory-based hot cache index layer, the method further includes: obtaining the address of ZooKeeper from the configuration file, establishing a connection with ZooKeeper, and obtaining all registered service processes; determining the location information of all service processes currently providing memory caching.

3. The method for reading a large number of small files of Hadoop with columnar storage format and hierarchical index optimization according to claim 1, characterized in that: Establishing an HBase original table based on the HBase-based picture storage strategy for storing picture small files and their attribute information includes: Establishing an HBase original table, where one column of the column family in the HBase original table stores the picture small file, and the picture small file is stored in this column in the form of binary data; Other columns of the column family in the HBase original table store various attribute information of the picture, and the attribute information includes small file type, size, creation time, and modification time.

4. The method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization according to claim 3, characterized in that: The HBase original table adopts a column-oriented storage model, and stores the data of the same column family continuously; When storing each column family, store the data in each row cell in the form of Key-Value to form several data blocks; Then save the data blocks to HFile, and finally save HFile to the background HDFS.

5. The method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization according to claim 1, characterized in that: After the index table established according to the non-row key attributes in the HBase original table, the method further includes: saving the index table in the HBase system to achieve persistent storage of index data for the Hbase original table; each index table is used to store the index of a certain non-row key attribute to be queried in the HBase original table; Define the format of the index table primary key for the non-row key attribute to be indexed in the HBase original table as follows: <original table index attribute name, original table index attribute value, original table row key>.

6. The method for reading a large number of small files of Hadoop with columnar storage format and hierarchical index optimization according to claim 1, wherein: When the query request includes multiple non-row key attributes at the same time, the index table established according to the non-row key attributes in the HBase original table includes: Construct a combined index table for multiple non-row key attribute columns; Through the established combined index table, convert the combined query into a query based on the index table primary key; Save the combined index table in the HBase system to achieve persistent storage of index data for the Hbase original table; each combined index table is used to store the index of a certain non-row key attribute to be queried in the HBase original table; Define the format of the combined index table primary key for the non-row key attribute to be indexed in the original table as follows: <original table index attribute name, original table index attribute value, original table row key>.

7. The method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization according to claim 1, characterized in that: The hot data cached in the hot cache index layer is obtained through the following steps: Based on the law that the access to network data satisfies the Pareto distribution, screen the user access log records on the HBase system, sort them according to the access volume, and select the top several users with the highest access volume as active users; Use the log-linear algorithm as the index hot data prediction model to calculate the hot values of the access data of each active user, and mark the picture small files with the top K% of the hot values as hot data; where K represents the hot threshold calculated according to the limit of the number of records that the cache space can accommodate; Cache the hot data in the hot cache index layer in a preset format.

8. The method for reading a large number of small files of Hadoop with columnar storage format and hierarchical index optimization according to claim 7, characterized in that: The calculation formula of the index hot data prediction model is as follows: ln N i = k(t)ln N i (t) + b(t) Where N t is the predicted total access volume of file i, that is, the hotspot value; N i (t) represents the access volume of file i within the observation time, and the length of the observation time is t; k(t) and b(t) are the relevant parameters of the linear relationship, and the optimal values are calculated using the linear regression method; the length of the observation time t represents the time difference between the access start time element of the record line in the user access log record and the time when the user access log record is collected.

9. The method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization according to claim 7, characterized in that: The format of the hot data cached in the hot cache index layer is as follows: Index primary key: <Original table index column name, Original table index column value> Index set: {<Original table primary key, {<Frequently accessed column name, Frequently accessed column value>}>}.

10. The method for reading a large number of small files in Hadoop with columnar storage format and hierarchical index optimization according to claim 1, characterized in that: The reading method further includes: when new pictures are written or existing pictures are modified or deleted, adopt an incremental update method to only update the indexes related to the changed pictures.