File merging method and device, electronic equipment and computer program product

By using data warehouse tools in HDFS to scan and evaluate the merging benefits of small files, combined with the automation characteristics of the Hive engine, the problem of small files cannot be automatically merged is solved, and the performance and management efficiency of the big data processing system is improved.

CN120492412APending Publication Date: 2025-08-15CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510464703.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art cannot realize automatic merging of small files at the HDFS global level, resulting in large performance overhead, complex management and inability to respond in real time.

Method used

By scanning the metadata of the distributed file system using data warehouse tools, evaluating and merging files with data sizes smaller than the preset threshold, combined with the powerful data processing capabilities and automation characteristics of the Hive engine, automatic merging of small files is achieved.

Benefits of technology

It significantly improves the performance and management efficiency of the big data processing system, realizes automatic merging of small files at the HDFS full domain level, and reduces the overhead of storage space and metadata operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492412A_ABST
    Figure CN120492412A_ABST
Patent Text Reader

Abstract

The invention discloses a file merging method and device, electronic equipment and a computer program product. The method comprises the steps that a data warehouse tool is used for scanning metadata corresponding to a plurality of files to be processed in the distributed file system, and the metadata is at least used for representing the data volume of the files to be processed; determining a to-be-merged file set according to a plurality of to-be-processed files of which the data volumes are smaller than a preset merging threshold value; evaluating the merging benefit after the plurality of to-be-merged files in the to-be-merged file set are merged; and under the condition that the merging benefit meets a preset benefit condition, merging a plurality of to-be-merged files in the to-be-merged file set, and under the condition that the saving amount of the storage space is greater than a preset space threshold value and the reduction amount of the metadata operation times is greater than a preset times threshold value, determining that the merging benefit meets the preset benefit condition. According to the method and the device, the technical problem that automatic merging of small files cannot be realized at the HDFS global level in a traditional mode is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data, and in particular to a file merging method, device, electronic device and computer program product. Background Art

[0002] In the field of big data processing, particularly data management and optimization involving Hadoop and Hive, a variety of existing technologies exist for handling small files in HDFS. The following are some examples of existing technologies:

[0003] 1. MapReduce-based small file merging solutions: Traditionally, MapReduce is used to handle small file merging in HDFS. These solutions usually involve writing custom MapReduce jobs to read and merge small files in batches, thereby reducing file system overhead and improving efficiency.

[0004] 2. Use Hadoop Archive (HAR): Hadoop Archive is a method for archiving multiple small files into a single file. It reduces the NameNode load and file system overhead by creating HAR files, but requires additional management and maintenance costs.

[0005] 3. Hive's ORC or Parquet file formats: ORC (Optimized Row Columnar) and Parquet are efficient columnar file formats supported by Hive. They are commonly used to store and process large amounts of data. However, they do not directly solve the problem of merging small files. Instead, they improve performance by optimizing data reading and storage.

[0006] 4. HDFS storage policy adjustment: Some methods indirectly improve the efficiency of the file system by adjusting the HDFS storage policy, such as increasing the block size or reducing the number of small files through file merging strategies.

[0007] 5. YARN resource manager tuning: YARN resource manager tuning can improve the resource utilization of the Hadoop cluster, thereby indirectly affecting the efficiency and optimization of small file processing.

[0008] Although these existing technologies are effective in processing small files in HDFS, they still have limitations such as performance overhead, operational complexity, or inability to be automated.

[0009] Therefore, in existing technical solutions, processing small files in HDFS usually involves the following technical challenges:

[0010] 1. Performance overhead and efficiency issues: Traditional MapReduce jobs or Hadoop Archive methods require a large number of I / O operations and data movement, which leads to slower processing speeds and reduced resource utilization. Especially on large-scale data sets, the efficiency of these methods may not be sufficient to meet real-time processing or high-throughput requirements.

[0011] 2. Complex management and maintenance: Using MapReduce to merge small files requires writing and maintaining a large number of custom jobs, which is a technical barrier for users unfamiliar with MapReduce programming. In addition, Hadoop Archive requires additional management work to maintain archive files and is not suitable for scenarios where small files are generated dynamically and in real time.

[0012] 3. Limitations of Automation: Most existing methods require manual triggering or periodic job execution to process small files, lacking automation capabilities. This increases operational complexity and management costs in large-scale and dynamically changing data environments.

[0013] Although existing technologies can process small files in some scenarios, they still have significant limitations in automation, real-time performance, and performance optimization.

[0014] Currently, no effective solution has been proposed to the problem that the above traditional method cannot automatically merge small files at the HDFS global level. Summary of the Invention

[0015] Embodiments of the present invention provide a file merging method, device, electronic device, and computer program product to at least solve the technical problem that traditional methods cannot achieve automatic merging of small files at the HDFS global level.

[0016] According to one aspect of an embodiment of the present invention, a file merging device is provided, comprising: using a data warehouse tool to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; determining a file set to be merged based on the multiple files to be processed whose data volume is smaller than a preset merging threshold, wherein the file set to be merged includes multiple files to be merged; evaluating a merging benefit after merging the multiple files to be merged in the file set to be merged, wherein the merging benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; merging the multiple files to be merged in the file set to be merged if the merging benefit meets a preset benefit condition, wherein if the amount of storage space saved is greater than a preset space threshold and the amount of reduction in the number of metadata operations is greater than a preset number threshold, determining that the merging benefit meets the preset benefit condition.

[0017] Optionally, using a data warehouse tool to scan the metadata corresponding to multiple files to be processed in a distributed file system includes: using the structured query language function of the data warehouse tool to scan the metadata corresponding to multiple files to be processed stored in the distributed file system; or monitoring the data stream of the distributed file system to obtain the metadata corresponding to multiple files to be processed in the data stream.

[0018] Optionally, after scanning metadata corresponding to a plurality of files to be processed, the apparatus further comprises: recording the metadata using a preset data structure, wherein the preset data structure comprises at least a hash table or a tree structure.

[0019] Optionally, the metadata is also used to indicate the file type of the file to be processed. After determining the file set to be merged based on the multiple files to be processed whose data volume is smaller than a preset merge threshold, the device further includes: setting a corresponding merge priority for each file to be merged in the file set to be merged based on the file type indicated by the metadata, wherein different file types correspond to different merge priorities.

[0020] Optionally, the metadata is also used to represent the access frequency of the files to be processed. After determining the set of files to be merged based on the multiple files to be processed whose data volume is smaller than a preset merge threshold, the device further includes: setting a corresponding merge priority for each file to be merged in the set of files to be merged based on the access frequency indicated by the metadata, wherein the access frequency is negatively correlated with the merge priority.

[0021] Optionally, evaluating the merging benefit after merging the multiple files to be merged in the file set to be merged includes: obtaining multiple preset merging strategies, wherein the preset merging strategy is used to indicate that all or part of the files to be merged in the file set to be merged are merged into a target file; evaluating the merging benefit corresponding to each preset merging strategy, wherein the preset merging strategy whose merging benefit meets the preset benefit condition is the target merging strategy, and the multiple files to be merged in the file set to be merged perform the merging operation according to the target merging strategy.

[0022] Optionally, when the merging benefit meets the preset benefit conditions, merging the multiple files to be merged in the file set to be merged includes: when the merging benefit meets the preset benefit conditions, copying the multiple files to be merged that need to be merged in the file set to be merged to a temporary directory; performing a merging operation on the files to be merged in the temporary directory to obtain a target file; and adding metadata of the target file in the distributed file system.

[0023] According to another aspect of an embodiment of the present invention, a file merging device is also provided, including: a scanning module, used to use a data warehouse tool to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; a determination module, used to determine a set of files to be merged based on the multiple files to be processed whose data volume is less than a preset merging threshold, wherein the set of files to be merged includes multiple files to be merged; an evaluation module, used to evaluate the merging benefit after merging the multiple files to be merged in the set of files to be merged, wherein the merging benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; a merging module, used to merge the multiple files to be merged in the set of files to be merged when the merging benefit meets the preset benefit condition, wherein when the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit condition.

[0024] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned file merging method through the computer program.

[0025] According to another aspect of an embodiment of the present invention, a computer program product is provided, including computer instructions, which implement the steps of the above-mentioned file merging method when executed by a processor.

[0026] In an embodiment of the present invention, a data warehouse tool is used to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; a file set to be merged is determined based on multiple files to be processed whose data volume is less than a preset merge threshold, wherein the file set to be merged includes multiple files to be merged; a merging benefit after merging the multiple files to be merged in the file set to be merged is evaluated, wherein the merging benefit is at least used to represent the amount of storage space saved after the merger and the amount of reduction in the number of metadata operations; when the merging benefit meets the preset benefit condition, the multiple files to be merged in the file set to be merged are merged, wherein in the storage space When the amount of space saved is greater than the preset space threshold and the amount of metadata operation reduction is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit conditions. By combining the powerful data processing capabilities and automation characteristics of the Hive engine, the problems of low small file processing efficiency, complex operations, and inability to respond in real time in the existing technology are solved, thereby significantly improving the performance and management efficiency of the big data processing system, achieving the purpose of significantly improving the performance and management efficiency of the big data processing system, and realizing the technical effect of automatic merging of small files at the HDFS global level, thereby solving the technical problem that the traditional method cannot achieve automatic merging of small files at the HDFS global level. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0028] Figure 1 is a flowchart of a file merging method according to an embodiment of the present invention;

[0029] Figure 2 Schematic diagram of a method for identifying and automatically merging global small files in HDFS based on the Hive engine according to an embodiment of the present invention;

[0030] Figure 3 is a schematic diagram of a file merging device according to an embodiment of the present invention;

[0031] Figure 4 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0034] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0035] HDFS, or Hadoop Distributed File System, is a key component of Hadoop, primarily used to store large amounts of data. It's a distributed file system designed to distribute data across a cluster of computers, providing high-throughput access to meet the demands of large-scale data processing. It's ideal for big data applications that span hundreds or even thousands of servers.

[0036] Hive is a data warehouse tool based on Hadoop. It maps structured data files into database tables and provides complete SQL query functionality. It can convert SQL statements into MapReduce tasks for execution, allowing Hadoop users to use the SQL-like Hive SQL language instead of MapReduce to process large amounts of data on Hadoop. Its primary function is to facilitate data access, improve the speed and efficiency of data processing, and make it easier for data analysts to perform statistics and analysis.

[0037] OIV, short for Offline Image Viewer, is a tool for viewing the contents of a Hadoop file system metadata image file (faimage). It can store the contents of the fsimage file in a specified file for easy reading.

[0038] MR stands for MapReduce, a programming model and algorithm for big data processing. It is used to decompose large-scale data sets into multiple small data blocks, and process data in parallel on a distributed computing cluster through the Map and Reduce stages to improve efficiency.

[0039] fsimage is a core component of Hadoop's Distributed File System (HDFS). Specifically, it is a persistent file used by the HDFS NameNode to store metadata. In HDFS, the NameNode manages the file system's namespace, specifically the metadata for files and directories, including file names, permissions, attributes, and operations such as file and directory creation and deletion.

[0040] Metadata refers to data used to describe, explain, locate, or otherwise help understand data. It can be considered as information about data, providing information about its content, quality, status, and other characteristics.

[0041] Apache Flink is an open-source stream processing framework for high-performance, complex processing on both unbounded and bounded data streams. Flink's key features include low-latency and high-throughput processing of streaming data, as well as powerful windowing and event-time processing capabilities, making it well-suited for real-time data analysis.

[0042] Spark Streaming is a module of Apache Spark that is used to process real-time data streams. It implements stream processing by splitting the data stream into a series of tiny batches that are processed at time intervals (such as seconds).

[0043] A big data PaaS (Platform as a Service) platform provides data processing, analysis, storage, and management services within a cloud computing architecture. It's a mid-tier service, located above IaaS (Infrastructure as a Service) and below SaaS (Software as a Service), offering users a service model that allows them to process and analyze big data without having to manage and maintain the underlying hardware and software environment.

[0044] According to an embodiment of the present invention, an embodiment of a file merging method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] Figure 1 is a flow chart of a file merging method according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0046] Step S102: Scan metadata corresponding to a plurality of files to be processed in a distributed file system using a data warehouse tool, wherein the metadata is used to at least indicate the data volume of the files to be processed;

[0047] Step S104, determining a set of files to be merged based on the plurality of files to be processed whose data volumes are smaller than a preset merging threshold, wherein the set of files to be merged includes the plurality of files to be merged;

[0048] Step S106, evaluating the merging benefit of the plurality of files to be merged in the file set to be merged, wherein the merging benefit is used to represent at least the amount of storage space saved and the amount of metadata operation times reduced after the merging;

[0049] Step S108, when the merging benefit meets the preset benefit conditions, merge the multiple files to be merged in the file set to be merged, wherein, when the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit conditions.

[0050] In an embodiment of the present invention, a data warehouse tool is used to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; a file set to be merged is determined based on multiple files to be processed whose data volume is less than a preset merge threshold, wherein the file set to be merged includes multiple files to be merged; a merging benefit after merging the multiple files to be merged in the file set to be merged is evaluated, wherein the merging benefit is at least used to represent the amount of storage space saved after the merger and the amount of reduction in the number of metadata operations; when the merging benefit meets the preset benefit condition, the multiple files to be merged in the file set to be merged are merged, wherein in the storage space When the amount of space saved is greater than the preset space threshold and the amount of metadata operation reduction is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit conditions. By combining the powerful data processing capabilities and automation characteristics of the Hive engine, the problems of low small file processing efficiency, complex operations, and inability to respond in real time in the existing technology are solved, thereby significantly improving the performance and management efficiency of the big data processing system, achieving the purpose of significantly improving the performance and management efficiency of the big data processing system, and realizing the technical effect of automatic merging of small files at the HDFS global level, thereby solving the technical problem that the traditional method cannot achieve automatic merging of small files at the HDFS global level.

[0051] In the above-mentioned embodiment of the present application, by evaluating the merging benefit after merging multiple files to be merged in the file set to be merged, and merging multiple files to be merged in the file set to be merged when the merging benefit meets the preset benefit conditions, the merging result of the files to be merged can have better performance and be more in line with the expected effect of file merging.

[0052] In the above step S102, the data warehouse tool may be a Hive engine having SQL query capabilities.

[0053] In the above step S102, the distributed file system may be an HDFS system.

[0054] In the above step S102, the data warehouse tool may scan the metadata according to a preset time period, or may scan the metadata when the total volume of updated files in the distributed file system reaches a preset volume threshold.

[0055] In the above step S102 , the metadata may further indicate the storage location of the file to be processed, and the metadata may enable acquisition of the file to be processed.

[0056] In the above step S102, the metadata can be expressed in the following form:

[0057]

[0058]

[0059] In the above step S104 , the file to be merged may be a small file, that is, a file to be processed whose data volume is smaller than a preset merging threshold.

[0060] It should be noted that the multiple files to be processed may have multiple file types, and therefore the files to be merged determined based on the files to be processed may also have multiple file types.

[0061] In the above step S104 , the set of files to be merged may include multiple files to be merged with different file types. When merging the files to be merged, the files to be merged with the same file type may be screened from the set of files to be merged for merging.

[0062] In the above step S104, there may be multiple file sets to be merged, and different file sets to be processed correspond to different file types. Each file set to be merged has multiple files to be merged with the same file type. When merging the files to be merged, the merging can be performed separately for each file set to be merged.

[0063] In the above step S106 , the step of evaluating the merging benefit may be performed before merging the files to be merged, or may be performed asynchronously during the merging process of the files to be merged.

[0064] In the above step S106, multiple files to be merged in the file set to be merged can be pre-merged, and an evaluation can be performed based on the pre-merger result. If the evaluation result is that the merger benefit meets the preset benefit conditions, the pre-merger result is used as the final merger result, thereby completing the merger of the merged files.

[0065] In the above step S108 , merging the multiple files to be merged in the file set to be merged includes: merging all the files to be merged in the file set to be merged into one target file; or selecting some of the files to be merged from the file set to be merged for merging.

[0066] Optionally, some of the files to be merged are selected from the set of files to be merged for merging. This can be done by selecting files to be merged of the same file type from the set of files to be merged, or by selecting multiple files to be merged that meet the priority merging conditions and merge them according to the merging priority set for the files to be merged.

[0067] Optionally, when merging multiple files to be merged in the file set to be merged, the merging process may be monitored and a report may be generated.

[0068] As an optional embodiment, using a data warehouse tool to scan the metadata corresponding to multiple files to be processed in a distributed file system includes: using the structured query language function of the data warehouse tool to scan the metadata corresponding to multiple files to be processed stored in the distributed file system; or monitoring the data stream of the distributed file system to obtain the metadata corresponding to multiple files to be processed in the data stream.

[0069] In the above-mentioned embodiment of the present application, the data warehouse tool can obtain the metadata corresponding to the files to be processed stored in the distributed file system, and can also perform indirect monitoring by querying and analyzing the data stored in the distributed file system, thereby realizing the monitoring of the data flow of the distributed file system, and obtaining the metadata corresponding to multiple files to be processed in the data flow, thereby realizing the acquisition of metadata corresponding to various forms of files to be processed in the distributed file system.

[0070] It should be noted that by monitoring the data stream of the distributed file system and obtaining metadata corresponding to multiple files to be processed in the data stream, small files in the data stream can be merged in real time, thereby supporting rapid response and dynamic adjustment.

[0071] As an optional embodiment, after scanning metadata corresponding to a plurality of files to be processed, the apparatus further includes: recording the metadata using a preset data structure, wherein the preset data structure includes at least: a hash table or a tree structure.

[0072] In the above embodiment of the present application, after obtaining the metadata of the file to be processed by scanning, a hash table or a tree structure can be used to store and manage the scanned file information to avoid repeated scanning and improve retrieval efficiency.

[0073] As an optional embodiment, metadata is also used to indicate the file type of the file to be processed. After determining the file set to be merged based on multiple files to be processed whose data volume is smaller than a preset merge threshold, the device also includes: setting a corresponding merge priority for each file to be merged in the file set to be merged based on the file type indicated by the metadata, wherein different file types have different corresponding merge priorities.

[0074] In the above embodiment of the present application, metadata is also used to indicate the file type of the to-be-processed file, and the to-be-merged files in the to-be-merged file set are screened from the to-be-processed files. Therefore, the to-be-processed files also have corresponding metadata, and the metadata can also indicate the file type of the to-be-processed files. In addition, to-be-processed files of different file types have different merging benefits. Therefore, based on the file type of the to-be-processed files, corresponding merging priorities are assigned to the to-be-processed files of different file types, and files are merged based on the merging priorities, so that to-be-processed files with high merging benefits can be given priority.

[0075] For example, the merging efficiency of text files is higher than that of compressed files and database files. Therefore, by setting the merge priority, you can give priority to merging text files rather than merging compressed files or database files.

[0076] As an optional embodiment, metadata is also used to indicate the access frequency of the files to be processed. After determining the set of files to be merged based on multiple files to be processed whose data volume is smaller than a preset merge threshold, the device further includes: setting a corresponding merge priority for each file to be merged in the set of files to be merged based on the access frequency indicated by the metadata, wherein the access frequency is negatively correlated with the merge priority.

[0077] In the above embodiment of the present application, metadata is also used to indicate the access frequency of the files to be processed, and the files to be merged in the set of files to be merged are screened from the files to be processed. Therefore, the files to be processed also have corresponding metadata, and the metadata can also indicate the access frequency of the files to be processed. In addition, a low access frequency indicates that the file to be merged is rarely accessed, and the merging benefit after merging the files to be merged is higher. Therefore, a corresponding merging priority is assigned to the files to be processed according to the access frequency of the files to be processed, and files are merged based on the merging priority, so that the files to be processed with high merging benefits can be given priority.

[0078] Optionally, according to the access frequency indicated by the metadata, setting a corresponding merge priority for each file to be merged in the set of files to be merged includes: determining a preset frequency interval that matches the access frequency, wherein there are multiple preset frequency intervals, and each preset frequency interval is pre-set with a corresponding merge priority; according to the preset frequency interval corresponding to the access frequency, setting a corresponding merge priority for the file to be merged corresponding to the access frequency.

[0079] Optionally, when setting the merge priority based on the access frequency and file type of the files to be processed, different merge priorities or priority intervals may be assigned based on the access frequency and file type. For example, the merge priority may include at least a first priority, a second priority, a third priority, and a fourth priority. The first priority or the second priority may be assigned to the files to be merged based on the file type, and the third priority or the fourth priority may be assigned to the files to be merged based on the access frequency.

[0080] As an optional embodiment, evaluating the merging benefit after merging multiple files to be merged in a file set to be merged includes: obtaining multiple preset merging strategies, wherein the preset merging strategy is used to indicate that all or part of the files to be merged in the file set to be merged are merged into a target file; evaluating the merging benefit corresponding to each preset merging strategy, wherein the preset merging strategy whose merging benefit meets the preset benefit condition is the target merging strategy, and the multiple files to be merged in the file set to be merged perform the merging operation according to the target merging strategy.

[0081] In the above-mentioned embodiment of the present application, a plurality of preset merge strategies are pre-configured. Different files to be merged in the file set to be merged can be selected for merging based on different preset merge strategies. By evaluating the merging benefit of each preset merge strategy, a preset merge strategy whose merging benefit meets the preset benefit conditions can be selected from a plurality of preset merge strategies as the target merge strategy, and a merge operation can be performed based on the target merge strategy, so that the merge result has better performance and is more in line with the expected effect of file merging.

[0082] Optionally, the preset merge strategy considers the physical locations of the files to be merged, and merges multiple files to be merged whose physical locations are dispersed into a target file whose physical locations are continuous, so as to minimize disk fragmentation.

[0083] As an optional embodiment, when the merging benefit meets the preset benefit conditions, merging multiple files to be merged in the file set to be merged includes: when the merging benefit meets the preset benefit conditions, copying the multiple files to be merged that need to be merged in the file set to be merged to a temporary directory; performing a merging operation on the files to be merged in the temporary directory to obtain a target file; and adding metadata of the target file in the distributed file system.

[0084] In the above embodiment of the present application, when merging multiple files to be merged in a file set to be merged, the multiple files to be merged can be copied to a temporary directory; and the merge operation is performed on the files to be merged in the temporary directory to obtain a target file. After the merge is completed, the metadata of the target file is added to the distributed file system to implement the update of the metadata. Based on the updated metadata, the operation of the files to be merged can be transferred to the target file for implementation.

[0085] Optionally, after the files to be merged in the temporary directory are merged to obtain the target file, the files to be merged in the temporary directory and metadata corresponding to the files to be merged may be deleted.

[0086] Optionally, the multiple files to be merged may be determined based on a target merge strategy.

[0087] Optionally, the above-mentioned merging operation on the multiple files to be merged in the file set to be merged may be performed based on a target merging strategy.

[0088] The present invention also provides a preferred embodiment, which provides an HDFS global small file identification and automatic merging method based on the Hive engine. The preferred embodiment aims to realize automatic monitoring and merging of global small files by utilizing the characteristics of the Hive engine, thereby improving processing efficiency, reducing management complexity, and supporting dynamic data generation and real-time processing requirements. The main problem solved is the technical problem that traditional methods cannot realize automatic identification and merging of small files at the HDFS global level. In contrast, the use of the Hive engine to realize global-level small file identification and automatic merging emphasizes automation and system integration, and has higher technical innovation and practicality.

[0089] As an optional embodiment, a method for global small file identification and automatic merging of HDFS based on the Hive engine is used to implement global small file identification and automatic merging, and specifically includes the following steps:

[0090] Step S21: data scanning and recognition.

[0091] Optionally, use the Hive engine to scan HDFS and identify existing small files.

[0092] In the above step S21, the metadata and SQL capabilities of the Hive engine are used to efficiently scan and identify small files.

[0093] Step S22: Automatic merging strategy generation.

[0094] Optionally, based on the scan results, an automatic merging strategy is generated.

[0095] In the above step S22, a rule engine or algorithm is involved to determine which small files can be merged into a larger file.

[0096] Step S23: merging the jobs.

[0097] Optionally, a merge job is automatically triggered to merge the identified small files into a larger file.

[0098] In the above step S23, a merge task is generated by the Hive engine and executed on the HDFS.

[0099] Step S24: monitoring and feedback.

[0100] Optionally, monitor the merge process, record the merge effect and provide feedback to the system. This can include logging, performance metric collection, etc.

[0101] The above-mentioned embodiments of the present application, by combining the powerful data processing capabilities and automation characteristics of the Hive engine, solve the problems existing in the prior art such as low efficiency in small file processing, complex operations, and inability to respond in real time, thereby significantly improving the performance and management efficiency of the big data processing system.

[0102] The HDFS global small file identification and automatic merging method based on the Hive engine provided in this application is suitable for scenarios that need to process a large number of dynamically generated small files, such as real-time data stream processing, log management systems, etc.

[0103] It should be noted that in big data processing platforms, especially data lake architectures based on the Hadoop ecosystem, the problem of large numbers of small files is often encountered. These small files may be dynamically generated by various data generation processes, including log files, daily and monthly data, etc. This application utilizes the metadata management and SQL query capabilities of the Hive engine to achieve global small file identification and automatic merging to optimize the storage and processing efficiency of the HDFS file system. The following are the main process steps of the technical solution and their logical relationship:

[0104] 1. Data scanning and identification. The execution entity is the Hive engine. The execution action is to scan the metadata of the HDFS file system through SQL queries to identify small files. The trigger condition is a scheduled task or when a certain amount of data is reached. The processing action is to mark and record the location, size, and related information of all small files. The processing result is to generate a small file list and identification report.

[0105] 2. Automatic merge strategy generation, the execution body is: automatic strategy generation module; the execution action is: formulate a merge strategy based on the scanning results; the trigger condition is: execute immediately after identifying small files; the processing action is: formulate a merge strategy and merge priority based on factors such as file size and storage location; the processing result is: generate a merge task queue or instruction.

[0106] 3. Merge job execution, the execution body is: Body: Hive engine and MapReduce job; Execution action: Automatically trigger MapReduce job to merge small files into larger files; Trigger condition: Automatic scheduling based on merge strategy and system load; Processing action: Merge selected small files into larger files and update the data storage structure in HDFS; Processing result: Reduce the number of small files and optimize storage and processing efficiency.

[0107] 4. Monitoring and feedback: The execution body is: monitoring and feedback module; the execution action is: real-time monitoring of the progress and effect of the merge job; the trigger condition is: after the merge job is completed or regular inspection; the processing action is: recording the merge log, collecting performance indicators and generating feedback reports; the processing result is: providing real-time system operation status and merge effect evaluation.

[0108] Figure 2 is a schematic diagram of a method for identifying and automatically merging HDFS global small files based on the Hive engine according to an embodiment of the present invention. Figure 2 As shown, multiple computer clusters in HDFS can regularly push fsimage files to the big data PaaS platform. The big data PaaS platform can then read the original text files through the OIV tool and use MapReduce jobs to decompose large-scale data sets into multiple small data blocks (such as small files) and store the images. The Hive engine can then be used to query, identify, and record small files in HDFS through metadata.

[0109] As an optional embodiment, a small file processing solution based on Hive metadata and SQL capabilities includes the following steps:

[0110] In step S211, the Hive engine identifies and records small files in HDFS through metadata query.

[0111] Step S212: The automated strategy generation module generates a merge strategy based on the file attributes.

[0112] Step S213: The MapReduce job performs file merging according to the strategy.

[0113] Step S214: The monitoring module tracks the merging process in real time and generates a report.

[0114] As an optional embodiment, for real-time data stream processing, the specific steps include:

[0115] Step S221: monitor the data stream in real time, and immediately identify and process small files.

[0116] Step S222: Dynamically adjust the merging strategy based on real-time demand.

[0117] Step S223: Use a stream processing engine (such as Apache Flink or Spark Streaming) to perform small file merging.

[0118] Step S224: Real-time feedback of processing effects to support rapid response and dynamic adjustment.

[0119] The above-mentioned embodiments of the present application reduce the number of small files in HDFS, reduce the management and processing overhead of the file system, thereby improving the efficiency of data processing and the performance of the overall system; automated small file identification and merging reduce the need for manual operations, and reduce the system's operation and maintenance costs and complexity; combined with real-time data stream processing, it can instantly identify and process dynamically generated small files, supporting the needs of real-time data analysis and application scenarios.

[0120] In general, the technical solution provided in this application, by combining the powerful functions and automation capabilities of the Hive engine, solves the problems of low efficiency and complex operation existing in traditional methods, and brings significant technical effects and practical application value to the big data processing platform.

[0121] As an optional embodiment, the global small file identification algorithm used in the HDFS global small file identification and automatic merging method based on the Hive engine includes:

[0122] 1.1. Use recursive or breadth-first search to traverse all files and directories in HDFS.

[0123] Optionally, for each file, get its size information.

[0124] 1.2. Definition of small files and setting of preset merge thresholds.

[0125] Optionally, a preset merging threshold for small files is defined. For example, the size of the preset merging threshold can be determined based on system configuration or performance testing.

[0126] Optionally, a common preset merging threshold may range from a few MB to dozens of MB, and is adjusted according to actual conditions.

[0127] 1.2. Real-time and efficiency.

[0128] Optionally, use an asynchronous task or a scheduled task to periodically scan the file system to ensure that small files are discovered in a timely manner.

[0129] Optionally, for newly written files, a notification mechanism of the file system can be combined to provide a quick response.

[0130] 1.3. Calculation of combined benefits.

[0131] Optionally, for the identified small files, a merging benefit after merging them is calculated, wherein the merging benefit may include reducing the number of metadata accesses, reducing storage space usage, improving read and write performance, etc.

[0132] 1.4. Data structure and storage.

[0133] Optionally, a suitable data structure (such as a hash table or a tree structure) is used to store and manage the scanned file information to avoid repeated scanning and improve retrieval efficiency.

[0134] As an optional example, the code logic of the above global small file identification algorithm is:

[0135]

[0136]

[0137] As an optional embodiment, the intelligent merging strategy used by the HDFS global small file identification and automatic merging method based on the Hive engine includes:

[0138] 2.1. Based on file size.

[0139] Optionally, a threshold for small files and merging (such as a preset merging threshold) is defined. For example, a small file is defined as a file with a size less than 100 MB.

[0140] 2.1. Based on file type.

[0141] Optionally, you can prioritize merging different types of files. For example, you can prioritize merging text files over compressed files or database files.

[0142] 2.1. Based on the metadata update amount.

[0143] Optionally, consider the frequency of file metadata updates. Files with frequently updated metadata may incur additional IO operations, and merging can reduce this overhead.

[0144] 2.1. Based on access frequency.

[0145] Optionally, analyze the file's access pattern. If a file is rarely accessed, consider merging it with other small files to reduce fragmentation on the storage device.

[0146] As an optional embodiment, the implementation of the merging strategy based on the HDFS global small file identification and automatic merging method of the Hive engine includes:

[0147] 3.1. The basic principle is: for each small file, evaluate whether it should be merged. First, check whether the file size is lower than a predetermined merging threshold (such as a preset merging threshold).

[0148] 3.2. Priority setting.

[0149] Optionally, you can set a merge priority based on file type and access pattern. For example, you can prioritize merging small files that are accessed infrequently.

[0150] 3.3. The merging process is: mark the small files that meet the merging conditions as pending merging status.

[0151] It should be noted that the physical location of the files should be considered when merging to minimize disk fragmentation.

[0152] 3.4. Calculation of combined benefits.

[0153] Optionally, the storage space savings (such as the amount of storage space saved) and the reduction in the number of metadata operations (such as the amount of reduction in the number of metadata operations) after the merger are calculated. These can be used as indicators for the merger decision.

[0154] As an optional example, the code logic of the above merge strategy is:

[0155]

[0156]

[0157] As an optional embodiment, a global small file identification and automatic merging method of HDFS based on the Hive engine can perform transparent merging, ensuring that the execution of the merge operation has no impact on Hive queries and does not affect the consistency and integrity of the data.

[0158] As an optional example, the merge operation may be performed during a non-high-load period or asynchronously.

[0159] Optionally, execution during non-high-load periods can ensure that the merge operation is performed when the Hive system load is low, such as during off-peak hours or when scheduled by a resource manager.

[0160] Optionally, the merge operation can be designed to be executed asynchronously to avoid blocking user queries or affecting system response time.

[0161] As an optional example, before executing a merge, analyze the currently executing and upcoming Hive queries, especially those involving files to be merged. Then, based on the query analysis results, formulate a merge plan to avoid merging the involved files during periods of active queries.

[0162] Optionally, for merging small files, it is necessary to ensure that the merge operation is transactional, that is, it either succeeds completely or does not perform any changes; this can be achieved through the transactional features of the file system or HDFS.

[0163] Optionally, before performing the merge operation, back up the relevant data to ensure timely recovery in case of unexpected situations.

[0164] As an optional example, the steps for implementing the merge operation include:

[0165] Step S1, file copying and merging: Copy the small files to be merged to a temporary directory, and perform the file merging operation in the temporary directory.

[0166] Step S2: Update metadata. Update the metadata in the file system or HDFS to ensure that the system can correctly identify and access the merged files.

[0167] Step S3: Delete the original file. After the merge is successful, the original small file is safely deleted to free up storage space.

[0168] As an optional example, the code logic for executing the merge strategy is as follows:

[0169]

[0170]

[0171] According to an embodiment of the present invention, a file merging device embodiment is also provided. It should be noted that the file merging device can be used to execute the file merging method in the embodiment of the present invention, and the file merging method in the embodiment of the present invention can be executed in the file merging device.

[0172] Figure 3 is a schematic diagram of a file merging device according to an embodiment of the present invention. Figure 3 As shown, the device may include: a scanning module 32, used to use a data warehouse tool to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; a determination module 34, used to determine a set of files to be merged based on multiple files to be processed whose data volume is less than a preset merge threshold, wherein the set of files to be merged includes multiple files to be merged; an evaluation module 36, used to evaluate the merging benefit after merging the multiple files to be merged in the set of files to be merged, wherein the merging benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; a merging module 38, used to merge the multiple files to be merged in the set of files to be merged when the merging benefit meets the preset benefit condition, wherein when the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit condition.

[0173] It should be noted that the scanning module 32 in this embodiment can be used to execute step S102 in the embodiment of the present application, the determining module 34 in this embodiment can be used to execute step S104 in the embodiment of the present application, the evaluating module 36 in this embodiment can be used to execute step S106 in the embodiment of the present application, and the merging module 38 in this embodiment can be used to execute step S108 in the embodiment of the present application. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments.

[0174] In an embodiment of the present invention, a data warehouse tool is used to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; a file set to be merged is determined based on multiple files to be processed whose data volume is less than a preset merge threshold, wherein the file set to be merged includes multiple files to be merged; a merging benefit after merging the multiple files to be merged in the file set to be merged is evaluated, wherein the merging benefit is at least used to represent the amount of storage space saved after the merger and the amount of reduction in the number of metadata operations; when the merging benefit meets the preset benefit condition, the multiple files to be merged in the file set to be merged are merged, wherein in the storage space When the amount of space saved is greater than the preset space threshold and the amount of metadata operation reduction is greater than the preset number threshold, it is determined that the merging benefit meets the preset benefit conditions. By combining the powerful data processing capabilities and automation characteristics of the Hive engine, the problems of low small file processing efficiency, complex operations, and inability to respond in real time in the existing technology are solved, thereby significantly improving the performance and management efficiency of the big data processing system, achieving the purpose of significantly improving the performance and management efficiency of the big data processing system, and realizing the technical effect of automatic merging of small files at the HDFS global level, thereby solving the technical problem that the traditional method cannot achieve automatic merging of small files at the HDFS global level.

[0175] As an optional embodiment, the scanning module includes: a scanning unit, which uses the structured query language function of the data warehouse tool to scan the metadata corresponding to multiple files to be processed stored in the distributed file system; or a detection unit, which monitors the data flow of the distributed file system and obtains the metadata corresponding to multiple files to be processed in the data flow.

[0176] As an optional embodiment, the apparatus further includes: a recording submodule, configured to record the metadata using a preset data structure after scanning metadata corresponding to a plurality of files to be processed, wherein the preset data structure includes at least: a hash table or a tree structure.

[0177] As an optional embodiment, metadata is also used to indicate the file type of the file to be processed, and the device also includes: a first setting unit, which is used to determine the file set to be merged based on multiple files to be processed whose data volume is smaller than a preset merge threshold, and then set a corresponding merge priority for each file to be merged in the file set to be merged based on the file type indicated by the metadata, wherein different file types have different corresponding merge priorities.

[0178] As an optional embodiment, the metadata is also used to indicate the access frequency of the files to be processed, and the device further includes: a second setting unit, which is used to, after determining the set of files to be merged based on multiple files to be processed whose data volume is smaller than a preset merge threshold, set a corresponding merge priority for each file to be merged in the set of files to be merged based on the access frequency indicated by the metadata, wherein the access frequency is negatively correlated with the merge priority.

[0179] As an optional embodiment, the evaluation module includes: an acquisition unit, used to obtain multiple preset merge strategies, wherein the preset merge strategy is used to indicate that all or part of the files to be merged in the file set to be merged are merged into a target file; an evaluation module, used to evaluate the merge benefit corresponding to each preset merge strategy, wherein the preset merge strategy whose merge benefit meets the preset benefit condition is the target merge strategy, and the multiple files to be merged in the file set to be merged perform the merge operation according to the target merge strategy.

[0180] As an optional embodiment, the merge module includes: a copy unit, which is used to copy multiple files to be merged in the file set to be merged to a temporary directory when the merge benefit meets the preset benefit conditions; a merge unit, which is used to perform a merge operation on the files to be merged in the temporary directory to obtain a target file; and an adding unit, which is used to add metadata of the target file in the distributed file system.

[0181] An embodiment of the present invention may provide an electronic device, which may be a computer terminal, and the computer terminal may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0182] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0183] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the file merging method: use a data warehouse tool to scan the metadata corresponding to multiple files to be processed in the distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; determine a file set to be merged based on multiple files to be processed whose data volume is less than a preset merging threshold, wherein the file set to be merged includes multiple files to be merged; evaluate the merging benefit after merging the multiple files to be merged in the file set to be merged, wherein the merging benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; if the merging benefit meets the preset benefit condition, merge the multiple files to be merged in the file set to be merged, wherein if the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, determine that the merging benefit meets the preset benefit condition.

[0184] Figure 4 is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 4 As shown, the computer terminal 40 may include: one or more (only one is shown in the figure) processors 42 and a memory 44.

[0185] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the file merging method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned file merging method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the terminal 40 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0186] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: use the data warehouse tool to scan the metadata corresponding to multiple files to be processed in the distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; determine the file set to be merged based on the multiple files to be processed whose data volume is less than a preset merge threshold, wherein the file set to be merged includes multiple files to be merged; evaluate the merger benefit after merging the multiple files to be merged in the file set to be merged, wherein the merger benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; if the merger benefit meets the preset benefit condition, merge the multiple files to be merged in the file set to be merged, wherein if the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, determine that the merger benefit meets the preset benefit condition.

[0187] Optionally, the processor may also execute the program code of the following steps: using the structured query language function of the data warehouse tool to scan the metadata corresponding to the multiple files to be processed stored in the distributed file system; or monitoring the data stream of the distributed file system to obtain the metadata corresponding to the multiple files to be processed in the data stream.

[0188] Optionally, the processor may further execute program code of the following steps: recording metadata using a preset data structure, wherein the preset data structure includes at least a hash table or a tree structure.

[0189] Optionally, the metadata is also used to indicate the file type of the file to be processed, and the above-mentioned processor can also execute the program code of the following steps: according to the file type indicated by the metadata, set a corresponding merging priority for each file to be merged in the set of files to be merged, wherein different file types correspond to different merging priorities.

[0190] Optionally, the metadata is also used to indicate the access frequency of the file to be processed, and the above-mentioned processor can also execute the program code of the following steps: according to the access frequency indicated by the metadata, set a corresponding merge priority for each file to be merged in the set of files to be merged, wherein the access frequency is negatively correlated with the merge priority.

[0191] Optionally, the processor may also execute the program code of the following steps: obtaining multiple preset merge strategies, wherein the preset merge strategy is used to indicate that all or part of the files to be merged in the file set to be merged are merged into a target file; evaluating the merge benefit corresponding to each preset merge strategy, wherein the preset merge strategy whose merge benefit meets the preset benefit condition is the target merge strategy, and the multiple files to be merged in the file set to be merged perform the merge operation according to the target merge strategy.

[0192] Optionally, the above-mentioned processor can also execute the program code of the following steps: when the merging benefit meets the preset benefit conditions, copy multiple files to be merged in the file set to be merged to a temporary directory; perform a merge operation on the files to be merged in the temporary directory to obtain the target file; and add metadata of the target file in the distributed file system.

[0193] It can be understood by those skilled in the art that Figure 4 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 4 It does not limit the structure of the above electronic device. For example, the computer terminal 40 may also include Figure 4 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 4 Different configurations shown.

[0194] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a computer program. The computer program can be stored in a non-volatile medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0195] The embodiment of the present invention further provides a non-volatile storage medium. Optionally, in this embodiment, the non-volatile storage medium can be used to store the program code executed by the file merging method provided in the embodiment.

[0196] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0197] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: using a data warehouse tool to scan metadata corresponding to multiple files to be processed in a distributed file system, wherein the metadata is at least used to represent the data volume of the files to be processed; determining a file set to be merged based on multiple files to be processed whose data volume is less than a preset merge threshold, wherein the file set to be merged includes multiple files to be merged; evaluating the merger benefit after merging the multiple files to be merged in the file set to be merged, wherein the merger benefit is at least used to represent the amount of storage space saved and the amount of reduction in the number of metadata operations after the merger; merging the multiple files to be merged in the file set to be merged when the merger benefit meets the preset benefit conditions, wherein when the amount of storage space saved is greater than the preset space threshold and the amount of reduction in the number of metadata operations is greater than the preset number threshold, determining that the merger benefit meets the preset benefit conditions.

[0198] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: using the structured query language function of the data warehouse tool to scan the metadata corresponding to multiple files to be processed stored in the distributed file system; or monitoring the data stream of the distributed file system to obtain the metadata corresponding to multiple files to be processed in the data stream.

[0199] Optionally, in this embodiment, the non-volatile storage medium is configured to store program codes for executing the following steps: recording metadata using a preset data structure, wherein the preset data structure at least includes: a hash table or a tree structure.

[0200] Optionally, in this embodiment, metadata is also used to indicate the file type of the file to be processed, and the non-volatile storage medium is configured to store program code for executing the following steps: based on the file type indicated by the metadata, setting a corresponding merging priority for each file to be merged in the set of files to be merged, wherein different file types have different corresponding merging priorities.

[0201] Optionally, in this embodiment, the metadata is also used to indicate the access frequency of the file to be processed, and the non-volatile storage medium is configured to store program code for executing the following steps: according to the access frequency indicated by the metadata, a corresponding merge priority is set for each file to be merged in the set of files to be merged, wherein the access frequency is negatively correlated with the merge priority.

[0202] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: obtaining multiple preset merge strategies, wherein the preset merge strategy is used to indicate that all or part of the files to be merged in the file set to be merged are merged into a target file; evaluating the merge benefit corresponding to each preset merge strategy, wherein the preset merge strategy whose merge benefit meets the preset benefit condition is the target merge strategy, and the multiple files to be merged in the file set to be merged perform the merge operation according to the target merge strategy.

[0203] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: when the merging benefit meets the preset benefit conditions, copying multiple files to be merged in the file set to be merged to a temporary directory; performing a merge operation on the files to be merged in the temporary directory to obtain a target file; and adding metadata of the target file in the distributed file system.

[0204] An embodiment of the present invention further provides a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the steps of the file merging method provided in the above embodiment are implemented.

[0205] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0206] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0207] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0208] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0209] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.

[0211] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A file merging method, characterized in that: include: Scanning metadata corresponding to a plurality of files to be processed in a distributed file system using a data warehouse tool, wherein the metadata is at least used to indicate a data volume of the files to be processed; Determining a set of files to be merged based on the plurality of files to be processed whose data volumes are smaller than a preset merging threshold, wherein the set of files to be merged includes a plurality of files to be merged; Evaluate a merging benefit after merging the plurality of files to be merged in the set of files to be merged, wherein the merging benefit is used to represent at least an amount of storage space saved and an amount of reduction in the number of metadata operations after the merging; When the merging benefit meets the preset benefit condition, multiple files to be merged in the set of files to be merged are merged, wherein, when the amount of storage space saved is greater than a preset space threshold and the amount of reduction in the number of metadata operations is greater than a preset number threshold, it is determined that the merging benefit meets the preset benefit condition.

2. The method according to claim 1, characterized in that Using data warehouse tools to scan the metadata corresponding to multiple files to be processed in the distributed file system includes: Using the structured query language function of the data warehouse tool, scanning the metadata corresponding to the plurality of files to be processed stored in the distributed file system; or The data stream of the distributed file system is monitored to obtain the metadata respectively corresponding to the plurality of files to be processed in the data stream.

3. The method according to claim 1, characterized in that After scanning metadata corresponding to the plurality of files to be processed, the method further includes: The metadata is recorded using a preset data structure, wherein the preset data structure at least includes: a hash table or a tree structure.

4. The method according to claim 1, wherein The metadata is further used to indicate the file type of the to-be-processed file. After determining a to-be-merged file set based on the plurality of to-be-processed files whose data volumes are smaller than a preset merging threshold, the method further includes: A corresponding merging priority is set for each of the files to be merged in the set of files to be merged according to the file type indicated by the metadata, wherein different file types correspond to different merging priorities.

5. The method according to claim 1, characterized in that The metadata is further used to indicate the access frequency of the to-be-processed files. After determining a set of files to be merged based on the plurality of to-be-processed files whose data volumes are smaller than a preset merging threshold, the method further includes: A corresponding merging priority is set for each of the to-be-merged files in the to-be-merged file set according to the access frequency indicated by the metadata, wherein the access frequency is negatively correlated with the merging priority.

6. The method according to claim 1, characterized in that Evaluating the merging benefit after merging the plurality of files to be merged in the set of files to be merged includes: Acquire multiple preset merging strategies, wherein the preset merging strategies are used to instruct to merge all or part of the files to be merged in the set of files to be merged into one target file; The merging benefit corresponding to each preset merging strategy is evaluated, wherein the preset merging strategy whose merging benefit meets the preset benefit condition is a target merging strategy, and the multiple files to be merged in the set of files to be merged are merged according to the target merging strategy.

7. The method according to claim 1, characterized in that When the merging benefit meets the preset benefit condition, merging the multiple files to be merged in the set of files to be merged includes: If the merging benefit meets the preset benefit condition, copying the plurality of files to be merged in the set of files to be merged to a temporary directory; Performing a merging operation on the files to be merged in the temporary directory to obtain a target file; Add metadata of the target file in the distributed file system.

8. A file merging device, characterized in that: include: A scanning module, configured to use a data warehouse tool to scan metadata corresponding to a plurality of files to be processed in a distributed file system, wherein the metadata is at least used to indicate the data volume of the files to be processed; a determination module, configured to determine a set of files to be merged based on the plurality of files to be processed whose data volumes are smaller than a preset merging threshold, wherein the set of files to be merged includes a plurality of files to be merged; An evaluation module, configured to evaluate a merging benefit after merging the plurality of files to be merged in the set of files to be merged, wherein the merging benefit is used to represent at least an amount of storage space saved and an amount of reduction in the number of metadata operations after the merging; A merging module is used to merge multiple files to be merged in the set of files to be merged when the merging benefit meets the preset benefit condition, wherein when the amount of storage space saved is greater than a preset space threshold and the amount of reduction in the number of metadata operations is greater than a preset number threshold, it is determined that the merging benefit meets the preset benefit condition.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the file merging method according to any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the file merging method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • File merging method and device, equipment and storage medium

    CN113127548A

  • File processing method and device, equipment, storage medium and computer program product

    CN117216009A

  • Small file merging method and device based on Flink engine and electronic equipment

    CN117632860A