Small file quick positioning and processing method and device of HDFS (Hadoop Distributed File System), electronic equipment and storage medium

By parsing the metadata generated by the fsimage file and positioning the location of the small file, the governance methods of small files in the HDFS distributed file system are solved, and the negative impact of small files on the HDFS system is improved, and the efficiency and stability of the system are improved.

CN120216476APending Publication Date: 2025-06-27中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371765.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The large number of small files in the HDFS distributed file system has serious negative impacts on the efficient and stable operation of the system, including excessive memory usage, degradation of performance, wasted storage space and difficulties in management and maintenance.

Method used

By parsing the new fsimage file generation into metadata and sending it to the database through a message queue, the metadata information is listened to to locate the location of the small file, and the governance of the data file and metadata file is performed based on the location.

Benefits of technology

It realizes rapid location and governance of small files, reduces access pressure to HDFS clusters, improves system efficiency and stability, and reduces management and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216476A_ABST
    Figure CN120216476A_ABST
Patent Text Reader

Abstract

The invention discloses a small file rapid positioning and processing method and device for an HDFS distributed file system, electronic equipment and a storage medium, and the method comprises the steps: when it is monitored that a new fsimage file is generated, analyzing the new fsimage file into metadata, and sending the metadata to a database through a message queue, the metadata comprises metadata information about files and metadata information about folders; monitoring the metadata information about the file and the metadata information about the folder in the database, and positioning to the position of a small file; and according to the position of the small file, respectively executing corresponding treatment means on the data file and the metadata file. According to the method and the device, a'closed-loop 'strategy for quickly finding and managing the small files in the HDFS is realized, the stability of cluster operation can be improved, meanwhile, full-automatic execution can be realized, and the efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of small file governance in the HDFS distributed file system, and particularly relates to a method, apparatus, electronic device, and storage medium for quickly locating and processing small files in the HDFS distributed file system. Background Art

[0002] HDFS (Hadoop Distributed File System), i.e., the distributed file system, plays an important basic support role in the big data processing framework and provides a reliable data storage service for upper-layer computing tasks.

[0003] Small files in the data warehouse established based on HDFS are still stored in the Hadoop cluster, which still poses a hazard to the HDFS system. For a data warehouse established based on HDFS, the large number of small files has a serious negative impact on the efficient and stable operation of the system. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, electronic device, and storage medium for quickly locating and processing small files in the HDFS distributed file system to optimize the processes of small file discovery, location, and governance.

[0005] Embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a method for quickly locating and processing small files in the HDFS distributed file system, where the method includes:

[0007] When a new fsimage file is monitored to be generated, parsing the new fsimage file into metadata and sending it to a database through a message queue, where the metadata includes metadata information about files and metadata information about folders;

[0008] Listening for the metadata information about files and the metadata information about folders in the database and locating the location where the small files are located;

[0009] According to the location where the small files are located, performing corresponding governance means on data files and metadata files respectively.

[0010] In some embodiments, the step of, when a new fsimage file is monitored to be generated, parsing the new fsimage file into metadata and sending it to a database through a message queue includes:

[0011] Deploying parsers on at least two main nodes Namenode of the HDFS distributed file system;

[0012] Through the parser, monitor in real time the folder where the fsimage cluster image file is generated;

[0013] When a new fsimage cluster image file is monitored, the parser determines whether the current Namenode is in the standby state;

[0014] If not, continue to monitor whether a new fsimage cluster image file is generated;

[0015] If so, read the fsimage cluster image file into memory to form a data set, and while traversing the data set, send the parsed metadata to the message queue to the database.

[0016] In some embodiments, the listening for the metadata information about the file and the metadata information about the folder in the database and locating the location of the small file includes:

[0017] Send the metadata information about the file to topic:file_info;

[0018] Send the metadata information about the folder to topic:dir_info;

[0019] Listen for the topic in the database and load the data in the topic into the database in real time.

[0020] In some embodiments, the listening for the metadata information about the file and the metadata information about the folder in the database and locating the location of the small file includes:

[0021] In the database data tables fsimage_file, fsimage_dir, and fsimage_commit, store file, folder, and image parsing time information respectively. The fsimage_file and the fsimage_dir respectively load the information of file_info and dir_info in the message queue. The table structure of the fsimage_commit includes fields of attributes, types, and meanings.

[0022] In some embodiments, the method further includes:

[0023] After completing the traversal of the content of the fsimage cluster image file, insert the information (fsimage_id, file_num, timestamp, 0) into fsimage_commit as a flag indicating the end of data parsing and sending.

[0024] In some embodiments, the method further includes:

[0025] By executing a first SQL query statement in the database, obtain the mirror ID of the latest fsimage cluster mirror file to avoid dirty reads of data;

[0026] By executing a second SQL query statement in the database, obtain the paths of files or folders whose accessed files and folders are smaller than a preset size.

[0027] In some embodiments, the corresponding governance means are respectively executed on the data file and the metadata file according to the location where the small file is located, including:

[0028] According to the location where the small file is located, locate the library name and table name where the small file is located;

[0029] According to the library name and table name where the small file is located, determine the number of metadata files and the number of data files;

[0030] If the number of metadata files is greater than the number of data files, execute the corresponding first governance means;

[0031] If the number of metadata files is less than the number of data files, execute the corresponding second governance means.

[0032] In a second aspect, an embodiment of the present application further provides a small file quick positioning and processing device for an HDFS distributed file system, where the device includes:

[0033] A discovery module, configured to parse the new fsimage file into metadata and send it to the database through a message queue when a new fsimage file is monitored, where the metadata includes metadata information about files and metadata information about folders;

[0034] A positioning module, configured to listen for the metadata information about files and the metadata information about folders in the database, and locate the location where the small file is located;

[0035] A governance module, configured to execute corresponding governance means on the data file and the metadata file respectively according to the location where the small file is located.

[0036] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the above method.

[0037] Fourthly, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device is caused to execute the above method.

[0038] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: When it is monitored that a new fsimage file is generated, the new fsimage file is parsed into metadata and sent to the database through a message queue. Since the metadata includes metadata information about files and metadata information about folders, by listening for the metadata information about files and the metadata information about folders in the database, the location where the small files are located can be located. Finally, according to the location where the small files are located, corresponding governance means are respectively executed on the data files and the metadata files. Through the present application, the discovery, location, and governance of small files are realized. During the discovery and location stage of small files, there is no need to access the cluster HDFS service. At the same time, concurrent technologies such as thread pools can be used to achieve efficient discovery and location. Through the above method, the discovery and location method of small files can facilitate the cluster administrator to locate the folders with more small files in the cluster and visually browse and query the metadata of the cluster, and can be used as an effective tool for the operation and maintenance of the Hadoop cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0040] FIG. 1(a) is a schematic structural diagram of the discovery, location, and governance modules of the small file quick location and processing method in the HDFS distributed file system according to the embodiment of the present application;

[0041] FIG. 1(b) is a schematic system architecture diagram of the small file quick location and processing method in the HDFS distributed file system according to the embodiment of the present application;

[0042] Figure 2 is a schematic flowchart of the small file quick location and processing method in the HDFS distributed file system according to the embodiment of the present application;

[0043] Figure 3 is a schematic diagram of the path of the data file in the small file quick location and processing method in the HDFS distributed file system according to the embodiment of the present application;

[0044] Figure 4 is a schematic diagram of the path of the metadata in the small file quick location and processing method in the HDFS distributed file system according to the embodiment of the present application;

[0045] Figure 5 This is a schematic diagram of the functions of the governance module in the method for quickly locating and processing small files in the HDFS distributed file system according to the embodiments of the present application;

[0046] Figure 6 This is a schematic diagram of the location process of the method for quickly locating and processing small files in the HDFS distributed file system according to the embodiments of the present application;

[0047] Figure 7 This is a schematic diagram of the governance process of the method for quickly locating and processing small files in the HDFS distributed file system according to the embodiments of the present application;

[0048] Figure 8 This is a schematic diagram of the structure of the device for quickly locating and processing small files in the HDFS distributed file system according to the embodiments of the present application;

[0049] Figure 9 This is a schematic diagram of the structure of an electronic device according to the embodiments of the present application. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0051] Technical terms involved in the embodiments of the present application:

[0052] HDFS: HDFS (Hadoop Distributed File System) is a distributed file system. It has the following main characteristics and advantages:

[0053] A. Distributed storage: Data is stored on multiple nodes, improving storage capacity and reliability.

[0054] B. High fault tolerance: It can automatically handle node failures and ensure data availability.

[0055] C. Large-scale data processing: It is suitable for processing massive amounts of data and can efficiently store and manage extremely large files.

[0056] D. Streaming data access: It supports the mode of writing once and reading multiple times, which is very suitable for data batch processing.

[0057] E. Strong scalability: The cluster scale can be easily expanded to meet the growing data requirements.

[0058] HDFS plays an important fundamental supporting role in the big data processing framework, providing a reliable data storage service for upper-layer computing tasks.

[0059] HDFS mainly consists of the following structures:

[0060] A. NameNode: It is the management node of the entire file system, responsible for managing the metadata of the file system, such as information about the names, locations, and sizes of files and directories.

[0061] B. DataNodes: Nodes that actually store data, and they store and retrieve data according to the instructions of the NameNode.

[0062] C. Data Block: Files are split into fixed-size data blocks for storage, which helps improve the read and write efficiency and fault tolerance of data.

[0063] D. Files and Directories: Users can create, delete, read, and write files and directories on HDFS.

[0064] The client obtains information such as file locations by interacting with the NameNode, and then directly performs data read and write operations with the DataNodes. This architecture enables HDFS to efficiently handle large-scale data storage and access requirements.

[0065] Data Warehouse: A data warehouse is a subject-oriented, integrated, relatively stable, and historical-change-reflecting data collection used to support management decision-making. Specifically, it has the following characteristics:

[0066] (1) Subject-oriented: Organize and classify data according to specific subjects, and data related to different subjects are concentrated together.

[0067] (2) Integrated: Extract, transform, and load data from multiple data sources to eliminate data inconsistencies and redundancies.

[0068] (3) Relatively stable: Mainly used for analysis, unlike the data in business systems that is updated frequently, the data is relatively stable.

[0069] (4) Reflecting historical changes: It stores a large amount of historical data for analyzing long-term trends and patterns. The data warehouse provides a basis for enterprise data analysis and decision support, helping decision-makers better understand the business situation, discover trends and patterns, and thus formulate more effective strategies and decisions.

[0070] Small Files: In HDFS, files with a size less than or equal to 32M are usually defined as small files. Too many small files will bring the following hazards to the HDFS system:

[0071] A. Excessive memory occupation: The NameNode in HDFS stores the metadata of the file system in memory, and each file, directory, and data block occupies approximately 150 bytes. If there are too many small files stored, it will occupy a large amount of memory and may even cause an out-of-memory error.

[0072] B. Performance degradation: When reading small files, it is necessary to jump between multiple DataNodes, which will increase the reading latency and overhead and reduce the system performance. In addition, for computing engines such as Hive or Spark, processing a large number of small files will also cause more tasks and resource consumption, affecting the computing speed.

[0073] C. Storage space waste: Since each small file will occupy an independent data block, even if the file content is very small, it will waste storage space. Especially when the number of small files is huge, this waste will be more obvious.

[0074] D. Difficult management and maintenance: A large number of small files increase the difficulty of system management and maintenance, and more time and effort are required to handle operations such as file creation, deletion, and movement.

[0075] FsImage, the metadata image of the Hadoop cluster master node:

[0076] In Hadoop, the NameNode is mainly responsible for managing metadata, including the directory structure of the file system, file and block information, etc. The NameNode stores the metadata in memory to improve data access speed. To achieve persistent storage of metadata, Hadoop uses two methods: FsImage (image file) and EditLog (edit log).

[0077] Specifically, when the NameNode modifies metadata, it will record the modification operations in the EditLog. The EditLog is a transaction file that records all modification operations on the metadata in chronological order. At the same time, the NameNode will regularly flush the metadata information in memory to the FsImage to create a new image file. This process is called Checkpoint. During the Checkpoint process, the NameNode will apply the modification operations in the EditLog to the FsImage to update the metadata information.

[0078] By combining the FsImage and EditLog, Hadoop can achieve persistent storage and fast recovery of metadata. When the NameNode crashes or restarts, the metadata can be restored by reading the information in the FsImage and EditLog, thus ensuring the consistency and availability of the file system.

[0079] Most data warehouses (such as Hive) and data lakes (such as Iceberg) are built relying on the HDFS system. Small files in the tables of the data warehouse are still stored on HDFS, which has a certain negative impact on the stable and efficient operation of HDFS.

[0080] However, the number of tables in the data warehouse is numerous, reaching the million level. Traversing each table not only has low efficiency and brings great pressure to the metadata service, which may cause the service to crash, but also there will be a phenomenon that metadata updates are not timely, and small files are not recorded by the metadata, so small files cannot be effectively discovered. However, too many small files have a great negative impact on both the cluster stability and the SQL batch running efficiency of the data warehouse, and cannot be ignored. However, there is no solution that can quickly locate and simultaneously manage small files in data lakes / warehouses built based on HDFS.

[0081] For example, the phenomenon that small files are not recorded by the metadata and thus cannot be effectively discovered: After the Hive SQL is killed during operation, the remaining cache directory:.hive-staging-20XX-XX-XX.... will not be recorded by the Hive metadata service (HiveMetastoreService).

[0082] In view of the above deficiencies, the embodiments of the present application provide a method for processing small files in the HDFS distributed file system, which performs automated processing of quickly locating and managing small files, reducing labor costs and improving efficiency. At the same time, the embodiments of the present application can also locate folders with more small files in the HDFS system and provide functions such as cluster metadata browsing, providing a reliable reference for the operation and maintenance of the HDFS system.

[0083] The following will, in conjunction with the accompanying drawings, detail the technical solutions provided by the embodiments of the present application.

[0084] As shown in Figure 1(a), it includes three major modules: discovery, location, and governance, which perform closed-loop governance on small files in the data warehouse based on HDFS.

[0085] The discovery module mainly uses the fsimage image in the parsing cluster to form a data set in memory and obtains information such as the size, path of the file, and the path of the folder by traversing the data set, as well as statistical information such as the total number of files, the number of files in the range of 0 - 8MB, 8 - 16MB, 16 - 32MB, 32 - 64MB, 64 - 128MB, 128 - 256MB, 256 - 512MB, and more than 512MB.

[0086] The positioning module mainly uses the storage principle of the data warehouse to parse the cluster path and obtain the library name, table name, and partition information, and locates the library tables with more small files through the sorting function of the OLAP database, such as Figure 3 and Figure 4 shown. Figure 3 and Figure 4 show the mapping relationship between the path information and the library name, table name, partition information, and data type (data file, metadata file). When parsing the mirror file, through this mapping relationship, relevant information such as the library name, table name, partition name, and file type corresponding to the file can be obtained.

[0087] The governance module mainly performs governance operations according to the file type after locating the library, table, partition, etc. information. These operations are mainly carried out by submitting spark SQL, Hive SQL, MapReduce tasks, etc. to the cluster.

[0088] Such as Figure 5 shown, taking the iceberg table as an example, if it is located that there are too many small files in the data folder of a certain table, methods such as deleting expired snapshots, rewriting data files by partition, and deleting orphan files can be used; if there are too many small files in the metadata folder of a certain table, methods such as deleting expired snapshots, rewriting metadata files, and reducing the number of metadata versions can be used.

[0089] For ordinary Hive tables stored in ORC format, the Hive SQL: "alter table db_name.tbl_name partition(partition_column=’XXX’)concatenate;" can be executed on the partitions with more small files. For small file merging, or adding small file merging parameters and rewriting the data files of the partition or even the whole table's data files.

[0090] In this way, through the closed-loop operations of discovery, positioning, and governance, the effective control of the number of small files in the data warehouse based on HDFS can be achieved.

[0091] As shown in Figure 1(b), it is the architecture of the data lake small file automated governance system designed according to the guiding ideology of "discovery", "positioning", and "governance". The parser deployed on two Namenode (master node) nodes monitors the folder where the cluster mirror file is generated in real time. If a new cluster mirror file appears, the parser first determines whether the Namenode is in the standby state. If not, it continues to monitor whether a new mirror file is generated. If so, it reads the mirror file into memory to form a data set, and sends the parsed metadata information to the message queue while traversing the data set.

[0092] Among them, the metadata information about files is sent to topic: file_info, and the metadata information about folders is sent to topic: dir_info. The OLAP database listens to the above two topics and loads the data in the two topics into the database in real time. The server in Figure 1(b) is responsible for analyzing the information in the OLAP database, locating the tables with many small files in the data warehouse, and by analyzing the folders with too many small files, taking corresponding measures to merge small files for data files and metadata files respectively, such as executing Spark SQL or Hive SQL, etc.

[0093] In specific practice, the definition of too many small files in a table can vary with the different sizes of cluster block blocks. Taking a cluster with a block size of 128MB as an example, a table with too many small files in the data warehouse can be defined as: the proportion of files with a size less than 64MB exceeds 20%. Users implementing the system proposed in the embodiments of the present application can adjust according to the situation of their own clusters.

[0094] Among them, the link for implementing the function of the "discovery" module is "parser" - "message queue (buffer layer)"; the link for implementing the function of the "location" module is "message queue (buffer layer)" - "OLAP database"; the link for implementing the function of the "governance" module is "OLAP database" - "server".

[0095] The embodiments of the present application provide a method for quickly locating and processing small files in an HDFS distributed file system, as Figure 2 shown, provides a schematic flowchart of the method for quickly locating and processing small files in the HDFS distributed file system in the embodiments of the present application. The method at least includes the following steps S210 to step S230:

[0096] Step S210, when a new fsimage file is monitored to be generated, parsing the new fsimage file into metadata and sending it to the database through a message queue. The metadata includes metadata information about files and metadata information about folders.

[0097] As shown in Figure 1(b), the parser is deployed on two main nodes of the cluster at the same time. When the main node state of the hadoop cluster is active, the parser on the main node is in the standby state and remains in the sleep state.

[0098] When the main node state is standby, the parser will change its own state to active, monitor the folder where the fsimage file is generated, and start parsing immediately once a new fsimage file is generated. This dual-node deployment ensures the high availability of the parser.

[0099] The fsimage is a key component in the Hadoop Distributed File System (HDFS) and is used to store a snapshot of the file system's metadata. The fsimage file contains the metadata information of the entire file system, including the hierarchical structure of files and directories, permissions, replication factor, and modification time. During the operation of HDFS, the fsimage file is continuously updated. Each time the NameNode starts, it loads the metadata from the fsimage file and then gradually applies the operations in the edits file to update the metadata.

[0100] Step S220: Listen for the metadata information about the files and the metadata information about the folders in the database, and locate the positions where the small files are located.

[0101] While parsing and traversing the fsimage file, the parser sends the file and folder information in the above two tables to the message queue in JSON format respectively. For example, use the corresponding topics: file_info and dir_info in the Kafka queue.

[0102] As shown in Figure 1(b), three tables need to be created in the OLAP database: fsimage_file, fsimage_dir, and fsimage_commit to store the relevant information about files, folders, and image parsing situations respectively. Among them, the fsimage_file and fsimage_dir need to stream and load the information in the file_info and dir_info in the message queue respectively. And the table structure of fsimage_commit is:

[0103]

[0104] Step S230: According to the positions where the small files are located, perform corresponding governance measures on the data files and metadata files respectively.

[0105] As shown in Figure 1(b), after the server accesses the OLAP database to locate the database name (db) and table name (tbl) with more small files, the service machine further analyzes and determines the number of metadata files and the number of data files. For different reasons, different measures are taken for governance.

[0106] Through the above method, using the fsimage file mirror information, the database name, table name, and whether the table needs to perform metadata file governance and data file governance are obtained by parsing the namenode metadata. There is no additional pressure on the cluster HDFS service during the stage of obtaining the relevant information of the table, and it can quickly locate the tables with more small files in the data warehouse.

[0107] Through the above method, an automated strategy for focusing on the discovery, location, and governance of small files in a data warehouse is provided. During the discovery and location phases of small files, without accessing the cluster HDFS service, concurrent technologies such as thread pools can be used to achieve efficient discovery and location. At the same time, the discovery and location methods of small files can facilitate cluster administrators to locate folders with a large number of small files in the cluster and perform intuitive browsing and querying of the cluster's metadata, becoming an effective tool for the operation and maintenance of the Hadoop cluster.

[0108] In an embodiment of the present application, when a new fsimage file is monitored to be generated, parsing the new fsimage file into metadata and sending it to the database through a message queue includes: deploying a parser on at least two primary nodes Namenode of the HDFS distributed file system; through the parser, monitoring in real time the folder where the fsimage cluster image file is generated; when a new fsimage cluster image file is monitored, the parser determines whether the current Namenode is in the standby state; if not, continue to monitor whether a new fsimage cluster image file is generated; if so, read the fsimage cluster image file into memory to form a data set, and while traversing the data set, send the parsed metadata to the database through the message queue.

[0109] When the primary node status is active, the parser on this node is in the standby state and remains in a sleeping state. When the primary node status is standby, the parser will change its own status to active, monitor the folder where the fsimage file is generated, and immediately start parsing once a new fsimage file is generated. Dual-node deployment is adopted to ensure the high availability of the parser.

[0110] The file information and types that the parser needs to obtain are as follows:

[0111]

[0112]

[0113] The folder information and types that the parser needs to obtain are as follows:

[0114]

[0115] Use the fsimage image in the parsing cluster to form a data set in memory, and obtain information such as the size and path of files, as well as statistical information such as the folder paths, total number of files, number of files in the range of 0 - 8MB, number of files in the range of 8 - 16MB, number of files in the range of 16 - 32MB, number of files in the range of 32 - 64MB, number of files in the range of 64 - 128MB, number of files in the range of 128 - 256MB, number of files in the range of 256 - 512MB, and number of files larger than 512MB in the folder by traversing the data set.

[0116] Through the above method, not only can small files in the HDFS system be located, but also small files in the data warehouse can be located.

[0117] In an embodiment of the present application, listening for the metadata information about files and the metadata information about folders in the database and locating the position of small files includes: sending the metadata information about files to topic: file_info; sending the metadata information about folders to topic: dir_info; listening for the topic in the database and loading the data in the topic into the database in real time.

[0118] While parsing and traversing the fsimage file, the parser sends the file and folder information in the above two tables to the corresponding topics: file_info and dir_info in a message queue such as Kafka, etc. in json format.

[0119] It should be noted that a topic refers to a theme or main information organization unit in the database. It usually represents a specific data set and contains all data related to a specific theme. By defining topics, relevant data can be stored centrally, so that the data needed can be found more quickly when querying data.

[0120] Through the above method, by adopting streaming updates, without the need to access the HDFS service, all metadata information of the cluster can be parsed and traversed, with high speed and efficiency. At the same time, it is deployed on two Namenode servers to achieve high availability.

[0121] Through the above method, a fully automated system that does not need to access the cluster service, but instead locates and takes governance measures for tables with a large number of small files in the HDFS-based data warehouse in near real time by automatically parsing the fsimage can greatly save labor costs.

[0122] In one embodiment of the present application, listening for the metadata information about files and the metadata information about folders in the database and locating the position where the small files are located includes: storing file, folder, and mirror parsing time information in the data tables fsimage_file, fsimage_dir, and fsimage_commit in the database respectively. The information in the fsimage_file and the fsimage_dir is respectively loaded into the information of file_info and dir_info in the message queue. The table structure of the fsimage_commit includes fields of attributes, types, and meanings.

[0123] Taking OLAP as an example for the database, three tables need to be established in the OLAP database: fsimage_file, fsimage_dir, and fsimage_commit to store information such as files, folders, and mirror parsing time respectively. Among them, the information in the fsimage_file and the fsimage_dir needs to be streamed and loaded into the information of file_info and dir_info in the message queue respectively. And the table structure of the fsimage_commit is as follows:

[0124]

[0125] In one embodiment of the present application, the method further includes: after completing the traversal of the content of the fsimage cluster mirror file, inserting the information (fsimage_id, file_num, timestamp, 0) into the fsimage_commit as a flag indicating the end of data parsing and sending.

[0126] After completing the traversal of the content of the fsimage file, the parser connects to the OLAP database and inserts the following information into the table fsimage_commit: (fsimage_id, file_num, timestamp, 0) as a flag indicating the end of data parsing and sending.

[0127] Through the above method, a system for quickly and efficiently parsing, displaying, and analyzing the metadata mirror of the Hadoop cluster is implemented in part of the architecture. Relying on the parser-message queue-OLAP database (including but not limited to Doris / Clickhouse / StarRocks, etc.) link, all metadata information in the namenode can be analyzed and displayed almost in real time. In actual use, after using technologies such as concurrency and thread pool, it only takes less than 3 minutes to parse and send the data in a 3GB mirror, and real-time monitoring can be achieved for small-scale clusters and near real-time monitoring can be achieved for large-scale clusters.

[0128] In one embodiment of the present application, the method further includes: obtaining the mirror ID of the latest fsimage cluster mirror file by executing a first SQL query statement in the database to avoid data dirty reads; obtaining the paths of files or folders whose accessed files and folders are smaller than a preset size by executing a second SQL query statement in the database.

[0129] By executing a first SQL query statement in the database, obtain the mirror ID of the latest fsimage cluster mirror file to avoid data dirty reads. Since data is continuously loaded in the OLAP database in a streaming manner, when selecting which mirror ID (fsimage_id) should be used to avoid data dirty reads when reading the mirror, the following SQL can be executed:

[0130] select fsimage_id from fsimage_commit where available = 1 order by timestamp desc limit 1; to obtain the latest fsimage mirror ID, thereby avoiding data dirty reads.

[0131] The main purpose of the first SQL query statement is to determine how many data records there are, and based on the information of these records, determine how many should be sent to the OLAP database, and query whether there is an update.

[0132] Obtaining the paths of files or folders whose accessed files and folders are smaller than a preset size by executing a second SQL query statement in the database means that the logic for locating the library tables with more small file problems is as follows:

[0133] select database_name,table_name from fsimage_dir where fsiamge_id = (select fsimage_id from fsimage_commit where available = 1 order by timestamp desc limit 1) and just_a_table order by file_num_0_8MB desc limit K;

[0134] Among them, K is the K tables with the most small files of 0 - 8MB. Of course, it can also be sorted in reverse order according to the number of files in other file size groups, such as 8 - 16MB.

[0135] The second SQL query statement is mainly used to determine the library and table locations where the small files are located.

[0136] Such as Figure 6As shown, preferably, taking OLAP as an example for the database, the OLAP database needs to start a scheduled task to maintain the available field and clean up historical expired data. The process is as follows:

[0137] Among them, to delete expired historical data, the following SQL needs to be executed:

[0138] delete from fsimage_file where fsimage_id in(select fsimage_id fromfsimage_commit where available=1order by timestamp desc offset K);

[0139] delete from fsimage_dir where fsimage_id in(select fsimage_id fromfsimage_commit where available=1order by timestamp desc offset K);

[0140] delete from fsimage_commit where fsimage_id in(select fsimage_id fromfsimage_commit where available=1order by timestamp desc offset K); In the SQL, K is the number of the latest images to be retained (k >= 1).

[0141] Among them, to maintain the available field, the following SQL needs to be executed:

[0142] Obtain the fsimage_id (id) that needs to be updated and the total number of files file_num (num1) corresponding to it

[0143] Execute the sql: select fsimage_id as id,file_num as num1 from fsimage_commitwhere available=0;

[0144] Query the number of files num2 that have been loaded for the image fsimage_id (id)

[0145] Execute the sql: select count(1)as num2 from fsimage_file where fsimage_id=id;

[0146] Compare whether num1 is equal to num2. If so, update the available field.

[0147] Execute the SQL: update fsimage_commit set available = 1 where fsimage_id = id;

[0148] If not, delete the expired historical data.

[0149] In an embodiment of the present application, the corresponding governance means are respectively executed on the data file and the metadata file according to the location where the small file is located, including: locating the library name and table name where the small file is located according to the location where the small file is located; determining the number of metadata files and the number of data files according to the library name and table name where the small file is located; if the number of metadata files is greater than the number of data files, execute the corresponding first governance means; if the number of metadata files is less than the number of data files, execute the corresponding second governance means.

[0150] After the server accesses the OLAP database to locate the library name (db) and table name (tbl) with more small files, the service machine further analyzes and determines the number of metadata files and the number of data files. Different means are taken for governance for different reasons.

[0151] As Figure 7 shown, a flowchart of the governance measures taken for the table (db.tbl) with serious small file problems is shown. Among them, the governance measures can be implemented in the form of submitting Spark SQL or Hive SQL to the cluster. Specifically, taking the iceberg table as an example, different APIs are called for governance. The corresponding relationship between the expressions in the embodiments of the present application and Spark SQL is as follows:

[0152]

[0153]

[0154] A queue can also be used to record the Spark SQL submitted to the cluster. Once it is executed, it dequeues, and the length of the queue is controlled according to the cluster resource status.

[0155] Specifically, taking the hive table as an example, since the hive table does not store metadata files in the data file, the hive table only needs to govern the data file, and different APIs are called for governance. The corresponding relationship between the expressions in the embodiments of the present application and hive SQL is as follows:

[0156]

[0157]

[0158] A queue can also be used to record the Hive SQL submitted to the cluster. Once the execution is completed, it will dequeue, and the length of the queue will be controlled according to the cluster resource status.

[0159] Through the above method, not only can the small files in the data warehouse based on HDFS be automatically managed, but also an effective tool can be provided for the cluster administrator to locate the cluster folders with too many small files, browse and analyze the cluster file information.

[0160] The embodiment of the present application also provides a small file quick location and processing device 800 for the HDFS distributed file system, as Figure 8 shown, which provides a structural schematic diagram in the embodiment of the present application. The small file quick location and processing device 800 for the HDFS distributed file system at least includes: a discovery module 810, a location module 820, and a governance module 830, where:

[0161] In an embodiment of the present application, the discovery module 810 is specifically configured to: when a new fsimage file is monitored to be generated, parse the new fsimage file into metadata and send it to the database through a message queue, where the metadata includes metadata information about files and metadata information about folders.

[0162] As shown in Figure 1(b), the parser is deployed on two main nodes of the cluster at the same time. When the main node status of the hadoop cluster is active, the parser on the main node is in the standby state and remains in the sleep state.

[0163] When the main node status is standby, the parser will change its own status to active, monitor the folder where the fsimage file is generated, and start parsing immediately once a new fsimage file is generated. This dual-node deployment ensures the high availability of the parser.

[0164] Fsimage is a key component in the Hadoop Distributed File System (HDFS) and is used to store a metadata snapshot of the file system. The fsimage file contains the metadata information of the entire file system, including the hierarchical structure of files and directories, permissions, replication factor, and modification time. During the operation of HDFS, the fsimage file will be continuously updated. Each time the NameNode starts, it will load the metadata from the fsimage file and then gradually apply the operations in the edits file to update the metadata.

[0165] In one embodiment of the present application, the positioning module 820 is specifically configured to: listen for the metadata information about the files and the metadata information about the folders in the database, and locate the positions where the small files are located.

[0166] While parsing and traversing the fsimage file, the parser sends the file and folder information in the above two tables to the message queue in the form of json respectively. For example, use the corresponding topics in the Kafka queue: file_info and dir_info.

[0167] As shown in Figure 1(b), three tables need to be established in the OLAP database: fsimage_file, fsimage_dir, and fsimage_commit to store the relevant information about files, folders, and image parsing conditions respectively. Among them, fsimage_file and fsimage_dir need to stream and load the information in the file_info and dir_info in the message queue respectively.

[0168] In one embodiment of the present application, the governance module 830 is specifically configured to: as shown in Figure 1(b), after the server accesses the OLAP database to locate the database name (db) and table name (tbl) with more small files, the service machine further analyzes and determines the number of metadata files and the number of data files. Different measures are taken for governance for different reasons.

[0169] In one embodiment of the present application, the discovery module 810 is further configured to

[0170] Deploy a parser on at least two master nodes Namenode of the HDFS distributed file system;

[0171] Through the parser, monitor the folder where the fsimage cluster image file is generated in real time;

[0172] When a new fsimage cluster image file is monitored, the parser determines whether the current Namenode is in the standby state;

[0173] If not, continue to monitor whether a new fsimage cluster image file is generated;

[0174] If so, read the fsimage cluster image file into the memory to form a data set, and send the parsed metadata to the database to the message queue while traversing the data set.

[0175] In one embodiment of the present application, the positioning module 820 is further configured to

[0176] Send the metadata information about the file to topic: file_info;

[0177] Send the metadata information about the folder to topic: dir_info;

[0178] Listen to the topic in the database and load the data in the topic into the database in real time.

[0179] In an embodiment of the present application, the positioning module 820 is further configured to

[0180] Store file, folder, and image parsing time information in the data tables fsimage_file, fsimage_dir, and fsimage_commit in the database respectively. The information in file_info and dir_info of the message queue is loaded into fsimage_file and fsimage_dir respectively. The table structure of fsimage_commit includes fields of attributes, types, and meanings.

[0181] In an embodiment of the present application, it further includes: a flag bit module, which is used to

[0182] After completing the traversal of the content of the fsimage cluster image file, insert the information (fsimage_id, file_num, timestamp, 0) into fsimage_commit as a flag indicating the end of data parsing and sending.

[0183] In an embodiment of the present application, it further includes: an SQL query module, which is used to

[0184] Obtain the latest fsimage cluster image file image ID by executing the first SQL query statement in the database to avoid dirty data reading;

[0185] Obtain the paths of files or folders whose accessed files and folders are smaller than a preset size by executing the second SQL query statement in the database.

[0186] In an embodiment of the present application, the governance module 830 is further configured to

[0187] Locate the library name and table name where the small file is located according to the location of the small file;

[0188] Determine the number of metadata files and the number of data files according to the library name and table name where the small file is located; if the number of metadata files is greater than the number of data files, execute the corresponding first governance measure; if the number of metadata files is less than the number of data files, execute the corresponding second governance measure.

[0189] It can be understood that the small file processing device of the above HDFS distributed file system can implement each step of the small file processing method of the HDFS distributed file system provided in the foregoing embodiments. The relevant explanations regarding the small file processing method of the HDFS distributed file system are applicable to the small file processing device of the HDFS distributed file system, and will not be elaborated here.

[0190] Figure 9 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 9 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include internal memory, such as high-speed random access memory (Random-Access Memory, RAM), and may also include non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0191] The processor, network interface, and memory can be interconnected through the internal bus. The internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 9 only a bidirectional arrow is used in

[0192] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory, and provide instructions and data to the processor.

[0193] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a small file processing device of the HDFS distributed file system at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0194] When it monitors the generation of a new fsimage file, it parses the new fsimage file into metadata and sends it to the database through a message queue. The metadata includes metadata information about files and metadata information about folders;

[0195] Monitor the metadata information about the file and the metadata information about the folder in the database, and locate the position where the small file is located;

[0196] According to the position where the small file is located, perform corresponding governance means on the data file and the metadata file respectively.

[0197] The above as in this application Figure 2 The method executed by the small file processing device of the HDFS distributed file system disclosed in the embodiments shown in this application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0198] The electronic device can also execute Figure 2 the method executed by the small file processing device of the HDFS distributed file system in, and implement the functions of the small file processing device of the HDFS distributed file system in Figure 2 the embodiments shown, which are not elaborated in the embodiments of this application.

[0199] An embodiment of the present application also provides a computer-readable storage medium storing one or more programs, where the one or more programs include instructions that, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 2 the method executed by the small file processing device of the HDFS distributed file system in the embodiment shown, and specifically used to execute:

[0200] When it is monitored that a new fsimage file is generated, parsing the new fsimage file into metadata and sending it to a database through a message queue, where the metadata includes metadata information about files and metadata information about folders;

[0201] Listening for the metadata information about files and the metadata information about folders in the database, and locating the positions where small files are located;

[0202] According to the positions where the small files are located, corresponding governance means are respectively executed on the data files and the metadata files.

[0203] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0204] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device that implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or multiple processes in the flowchart and / or one block or multiple blocks in the block diagram.

[0207] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0208] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0209] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0210] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, commodity or device including the said element.

[0211] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0212] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for quickly locating and processing small files in an HDFS distributed file system, wherein: The method comprises: When a new fsimage file is generated, the new fsimage file is parsed into metadata and sent to a database through a message queue, wherein the metadata includes metadata information about the file and metadata information about the folder; Monitoring the metadata information about the file and the metadata information about the folder in the database, and locating the location of the small file; According to the location of the small files, corresponding management measures are respectively executed on the data files and metadata files.

2. The method of claim 1, wherein: When monitoring the generation of a new fsimage file, parsing the new fsimage file into metadata and sending the metadata to the database through a message queue includes: Deploy the parser on at least two master nodes Namenode of the HDFS distributed file system; Through the parser, the folders generated by the fsimage cluster image files are monitored in real time; When a new fsimage cluster image file is detected, the parser determines whether the current Namenode is in the standby state; If not, continue to monitor whether the new fsimage cluster image file is generated; If so, the fsimage cluster image file is read into the memory to form a data set, and while traversing the data set, the parsed metadata is sent to the message queue to the database.

3. The method of claim 2, wherein: The step of monitoring the metadata information about the file and the metadata information about the folder in the database and locating the location of the small file includes: Send the metadata information about the file to topic:file_info; Send the metadata information about the folder to topic:dir_info; The topic is monitored in the database, and the data in the topic is loaded into the database in real time.

4. The method of claim 1, wherein: The step of monitoring the metadata information about the file and the metadata information about the folder in the database and locating the location of the small file includes: The data tables fsimage_file, fsimage_dir, and fsimage_commit in the database respectively store file, folder, and image parsing time information. The fsimage_file and the fsimage_dir respectively load the information of file_info and dir_info in the message queue. The table structure of the fsimage_commit includes fields of attributes, types, and meanings.

5. The method according to claim 4, further comprising: After completing the traversal of the fsimage cluster image file content, insert (fsimage_id, file_num, timestamp, 0) information into fsimage_commit as a sign of the end of data parsing and sending.

6. The method according to claim 5, further comprising: By executing the first SQL query statement in the database, the latest fsimage cluster image file image ID is obtained to avoid data dirty reading; By executing the second SQL query statement in the database, the path of accessing files or folders whose size is smaller than a preset size is obtained.

7. The method of claim 1, wherein: The performing corresponding management measures on the data file and the metadata file respectively according to the location of the small file includes: According to the location of the small file, locate the library name and table name where the small file is located; Determine the number of metadata files and the number of data files according to the library name and table name where the small files are located; If the number of metadata files is greater than the number of data files, executing the corresponding first governance measure; If the number of metadata files is less than the number of data files, the corresponding second governance measure is executed.

8. A device for quickly locating and processing small files in an HDFS distributed file system, wherein: The device comprises: A discovery module, for parsing the new fsimage file into metadata and sending the metadata to the database through a message queue when a new fsimage file is monitored to be generated, wherein the metadata includes metadata information about the file and metadata information about the folder; A positioning module, used for monitoring the metadata information about the file and the metadata information about the folder in the database, and locating the location of the small file; The management module is used to execute corresponding management measures on the data files and metadata files respectively according to the locations of the small files.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.