File processing method and device and computer readable storage medium

By building a Hive governance configuration table and grouping Spark computing tasks, small files of Hive partition tables and non-partition tables are merged, which solves the problem of small files generated by Hive's high-frequency data writing operations and improves storage and management efficiency.

CN120687412APending Publication Date: 2025-09-23CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510731554.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Hive generates a large number of small files during high-frequency data writing operations, resulting in low storage efficiency, increased metadata management pressure, and impact on file system performance and scalability. Existing technologies make it difficult to effectively coordinate and manage the small file problem.

Method used

Build a Hive governance configuration table, determine the merge subtasks and group them into Spark computing tasks, use the Spark computing task collection for serial processing, merge small files in partitioned and non-partitioned tables, merge files using the preset algorithm of the ORC engine, and update metadata after the merge is successful.

Benefits of technology

It achieves the rational allocation of computing resources, reduces data storage redundancy, lowers storage costs, avoids repeated calculations, and improves management efficiency and cluster resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687412A_ABST
    Figure CN120687412A_ABST
Patent Text Reader

Abstract

The invention discloses a file processing method and device and a computer readable storage medium, and relates to the technical field of data processing.The file processing method comprises the steps that a Hive governance configuration table is constructed, and the Hive governance configuration table comprises governance strategy parameters of a partition table and a non-partition table; determining at least one merging sub-task according to the governance strategy parameter, wherein the merging sub-task corresponds to the partition table directory to be merged or the non-partition table directory to be merged; determining the number of Spark calculation tasks corresponding to at least one merged subtask; the merged sub-tasks are grouped according to the number of the Spark calculation tasks, a corresponding Spark calculation task set is constructed according to the grouped merged sub-tasks, and each Spark calculation task is used for serial processing of the allocated merged sub-tasks. According to the data processing method and device, repeated calculation on repeated data in the data processing process is avoided, and the management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a file processing method, device, and computer-readable storage medium. Background Art

[0002] In big data applications, the data warehouse tool Hive is widely used to execute high-concurrency daily batch tasks, and occupies a key position in enterprise data processing and analysis. However, with the continuous growth of data volume and increasing business complexity, Hive has exposed a series of significant problems in high-frequency data write operations. These problems mainly stem from the large number of small files generated when Hive writes to HDFS (Hadoop Distributed File System). The size of these small files is far smaller than the block size configured for HDFS, resulting in low storage efficiency. Regardless of size, each file requires NameNode memory to store its metadata. Therefore, the surge in small files will excessively occupy NameNode memory, increase the pressure on metadata management, and thus affect the performance and scalability of the entire file system.

[0003] Currently, single-job merging solutions based on Hive job parameters are being attempted to address the small file problem. This solution configures Hive parameters to trigger small file merging at the end of a single job. However, this approach cannot achieve batch management across tables or multiple partitions, making it difficult to effectively address the small file problem in large-scale clusters. Furthermore, it lacks comprehensive management and efficient means to address small file issues across the entire data warehouse. Summary of the Invention

[0004] The main purpose of this application is to provide a file processing method, device and computer-readable storage medium, aiming to solve the technical problem of how to comprehensively manage small files.

[0005] To achieve the above objectives, the present application proposes a file processing method, which includes:

[0006] Constructing a Hive governance configuration table, which includes governance policy parameters for partitioned tables and non-partitioned tables;

[0007] Determine at least one merging subtask according to the governance policy parameters, where the merging subtask corresponds to the partition table directory or the non-partition table directory to be merged;

[0008] Determine the number of Spark computing tasks corresponding to at least one of the merged subtasks;

[0009] The merged subtasks are grouped according to the number of the Spark computing tasks, and a corresponding Spark computing task set is constructed according to the grouped merged subtasks. Each Spark computing task is used to serially process the assigned merged subtasks.

[0010] In one embodiment, the governance policy parameter includes a concurrency configuration threshold, and the step of determining the number of Spark computing tasks corresponding to at least one of the merged subtasks includes:

[0011] Determining a first configuration concurrency number according to the concurrency configuration threshold and the total number of the merged subtasks;

[0012] The number of Spark computing tasks is determined according to the first configured concurrency number.

[0013] In one embodiment, the step of determining the number of Spark computing tasks according to the first configured concurrency number includes:

[0014] Determine the second configuration concurrency based on the number of available cores in the cluster;

[0015] Determine the minimum configured concurrency number between the first configured concurrency number and the second configured concurrency number;

[0016] The number of Spark computing tasks is determined based on the minimum configured concurrency.

[0017] In one embodiment, the governance policy parameters include a table name, a partition importance flag, and a non-partition flag, and the step of determining at least one merge subtask based on the governance policy parameters includes:

[0018] Determine the partition table directory to be merged according to the table name and the partition importance mark;

[0019] Determine a non-partition table directory to be merged according to the table name and the non-partition mark;

[0020] At least one merging subtask is generated according to the partition table directory to be merged or the non-partition table directory.

[0021] In one embodiment, the step of generating at least one merging subtask according to the partition table directory to be merged or the non-partition table directory includes:

[0022] Obtain directory information of the partition table directory or the non-partition table directory to be merged, wherein the directory information includes a task identifier, a table name, and an HDFS partition path; wherein the HDFS partition path includes a Hive root directory and a table partition path;

[0023] The merging subtask is determined according to the directory information.

[0024] In one embodiment, after the step of constructing a corresponding Spark computing task set according to the grouped merged subtasks, the method further includes:

[0025] The Spark computing task set is distributed to an executor, so that the executor calls a preset algorithm of the ORC engine to merge the target files in the partition table directory or the non-partition table directory for each of the merge subtasks, and the size of the target file is smaller than a preset threshold.

[0026] In one embodiment, after the step of distributing the Spark computing task set to the executor, the method further includes:

[0027] During the merging process, a temporary directory is created under the target directory, and the temporary directory is used to store the merging results;

[0028] After the merge is successful, the temporary directory replaces the partition table directory or the non-partition table directory.

[0029] In one embodiment, after the step of distributing the Spark computing task set to the executor, the method further includes:

[0030] Monitor the merging progress and merging results of each Spark computing task in real time based on the accumulator;

[0031] Determining the partition table or the non-partition table that has been successfully merged according to the merge progress and the merge result;

[0032] A metadata repair operation is performed on the successfully merged partition table or the non-partition table.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes a file processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the file processing method described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by the processor, the steps of the file processing method described above are implemented.

[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the file processing method described above are implemented.

[0036] One or more technical solutions proposed in this application have at least the following technical effects:

[0037] The number of corresponding Spark computing tasks is determined based on the merged subtasks, and the merged subtasks are grouped to form a corresponding Spark computing task set, achieving the optimal allocation and utilization of computing resources. Each Spark computing task can focus on serially processing the assigned merged subtasks, avoiding excessive or insufficient resource consumption and improving cluster resource utilization. The merged subtasks can effectively consolidate and clean up duplicate or redundant data in partitioned and non-partitioned tables, reducing data storage redundancy and storage costs. It also avoids repeated computation of duplicate data during data processing, improving management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0040] Figure 1 A flowchart of the first embodiment of the method for processing application documents is provided;

[0041] Figure 2 A brief flowchart of the first embodiment of the method for processing application documents is provided;

[0042] Figure 3 A brief flowchart of the first embodiment of the method for processing application documents is provided;

[0043] Figure 4 A brief flowchart of the first embodiment of the method for processing application documents is provided;

[0044] Figure 5 A flowchart of the second embodiment of the method for processing application documents is provided;

[0045] Figure 6 A flowchart of the third embodiment of the method for processing application documents is provided;

[0046] Figure 7 A brief flowchart of Example 3 of the method for processing application documents is provided;

[0047] Figure 8A brief flowchart of Example 3 of the method for processing application documents is provided;

[0048] Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the file processing method in the embodiment of the present application.

[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0051] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0052] In big data applications, the data warehouse tool Hive is widely used to execute high-concurrency daily batch tasks, and occupies a key position in enterprise data processing and analysis. However, with the continuous growth of data volume and increasing business complexity, Hive has exposed a series of significant problems in high-frequency data write operations. These problems mainly stem from the large number of small files generated when Hive writes to HDFS (Hadoop Distributed File System). The size of these small files is far smaller than the block size configured for HDFS, resulting in low storage efficiency. Regardless of size, each file requires NameNode memory to store its metadata. Therefore, the surge in small files will excessively occupy NameNode memory, increase the pressure on metadata management, and thus affect the performance and scalability of the entire file system.

[0053] Currently, single-job merging solutions based on Hive job parameters are being attempted to address the small file problem. This solution configures Hive parameters to trigger small file merging at the end of a single job. However, this approach cannot achieve batch management across tables or multiple partitions, making it difficult to effectively address the small file problem in large-scale clusters. Furthermore, it lacks comprehensive management and efficient means to address small file issues across the entire data warehouse.

[0054] The main solution of the embodiment of the present application is: constructing a Hive governance configuration table, which includes governance policy parameters for partitioned tables and non-partitioned tables; determining at least one merge subtask based on the governance policy parameters, where the merge subtask corresponds to a partitioned table directory or a non-partitioned table directory to be merged; determining the number of Spark computing tasks corresponding to the at least one merge subtask; grouping the merge subtasks according to the number of Spark computing tasks, and constructing a corresponding Spark computing task set based on the grouped merge subtasks, where each Spark computing task is used to serially process the assigned merge subtasks.

[0055] In this embodiment, for ease of description, the following description is made with the file processing device as the execution subject.

[0056] The present application provides a solution that determines the number of corresponding Spark computing tasks based on the merged subtasks, and groups the merged subtasks to construct the corresponding Spark computing task set, thereby achieving the rational allocation and utilization of computing resources. Each Spark computing task can focus on serial processing of the assigned merged subtasks, avoiding excessive waste or insufficiency of resources and improving the utilization of cluster resources. For duplicate or redundant data in partitioned tables and non-partitioned tables, the processing of merged subtasks can be effectively integrated and cleaned up, reducing the redundancy of data storage and reducing storage costs. At the same time, it also avoids repeated calculation of duplicate data during data processing and improves management efficiency.

[0057] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or file processing device capable of implementing the above functions. This embodiment and the following embodiments will be described below using a file processing device as an example.

[0058] Based on this, the present application embodiment provides a file processing method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the document processing method of this application.

[0059] In this embodiment, the file processing method includes steps S10 to S40:

[0060] Step S10: construct a Hive governance configuration table, where the Hive governance configuration table includes governance policy parameters for partitioned tables and non-partitioned tables.

[0061] In this embodiment, a Hive governance configuration table is constructed on HDFS to unify governance policy parameters, implement governance operations, and coordinate partition cleanup and small file management for multiple tables.

[0062] It's important to note that HDFS is the core distributed storage component of the Apache Hadoop ecosystem. Designed for large-scale data storage, it offers high throughput, high fault tolerance, and horizontal scalability. HDFS utilizes a master-slave architecture, where the master node manages file system metadata, such as the directory tree and file block locations, while slave nodes store the actual data blocks and regularly report their status to the NameNode. HDFS is suitable for massive data batch processing scenarios, such as log analysis and data warehousing. However, an excessive number of small files can significantly increase the NameNode's memory pressure.

[0063] Hive is a data warehouse infrastructure built on HDFS, enabling large-scale data query and analysis through HQL (Hive Query Language). ORC (Optimized Row Columnar) is Hive's high-performance columnar storage format. ORC file structures utilize columnar storage, offering high compression ratios and excellent compatibility with Hive. This storage format improves Hive table query performance while consuming less storage space. Hive data stored in ORC format can be converted to ORC using HQL.

[0064] The small file problem exists in both HDFS and Spark. It manifests primarily in file sizes being significantly smaller than the block size, leading to wasted resources and reduced performance. In HDFS, small files increase the NameNode's memory load, impacting performance. In Spark, processing large numbers of small files increases task scheduling overhead and reduces efficiency.

[0065] Spark Dataset combines the elastic distribution characteristics of RDD (Resilient Distributed Dataset) with the structured processing capabilities of DataFrame. Spark Dataset provides type safety and efficient query optimization, making it suitable for complex data analysis tasks. DataFrame is a distributed, structured dataset widely used in distributed computing frameworks.

[0066] In this embodiment, the Hive governance configuration table is used to manage small files, wherein the small files in the partitioned table are located in the specific partition directory of the partitioned table and are generated by the daily new data of the table partitioned by time or region. The small files of the non-partitioned table are located in the root directory of the non-partitioned table and are generated by the full table or the dimension table with low frequency updates. Optionally, the partitioned table is located in the partition directory corresponding to the root directory, and the non-partitioned table is located in the root directory. By constructing the Hive governance configuration table, the governance policy parameters of the partitioned table and the non-partitioned table are uniformly managed, making the data governance work more systematic and standardized, avoiding the problems of inconsistent policies and chaotic management that may arise when different tables are governed separately, and improving the efficiency and standardization of data management.

[0067] The path structure for small files in a partitioned table is: Hive root directory / table name / partition key = value / file name.orc. Small files in partitions must be scanned partition by partition. For example, 365 date partitions x 100 city partitions = 36,500 directories. Partition-level metadata is also restored after merging. For example, the path structure for small files in a partitioned table is / user / hive / warehouse / order_db / orders / dt=20240528 / city=beijing / 000000_0.orc.

[0068] The path structure for small files in non-partitioned tables is: Hive root directory / table name / file name.orc. For example, / user / hive / warehouse / user_db / profile / user_info.orc. All small files in non-partitioned tables are located in a single directory, eliminating the need to traverse sub-directories. After merging, the table metadata must be refreshed.

[0069] Optionally, the governance policy parameters include table name, partition importance tag, concurrency configuration threshold, partition tag, non-partition tag, etc.

[0070] Optionally, redundant partition tables are pre-processed on the files under the partition directory according to the management policy parameters, and valid partition tables are retained as small files of the partition tables to be merged.

[0071] In one embodiment, referring to Figure 2 The Hive governance configuration table is used to uniformly manage the configuration of governance tables. Governed tables can be partitioned or non-partitioned. The governance strategy is to use a single job to archive important partitions. Important partitions are governance tables manually marked with the "important" flag "Y" in the configuration table. Redundant partitions are then cleaned up, and small files in the partitioned or non-partitioned tables corresponding to the Hive governance configuration table are merged.

[0072] In one embodiment, if Figure 3During the data preparation phase, partition cleanup and retention strategies are implemented. The Spark job first performs a series of validity checks on the small files to be merged. If they pass, the pending partitions of the partitioned table and the non-partitioned table are merged into a list, which serves as the data preparation stage for the small file merge list. The accumulator facilitates the statistics of various merge scenario indicators. After registration, a listening thread is enabled to monitor the execution progress of the executor in real time. At the end of the merge, the accumulator is responsible for listing the partitions awaiting metadata repair.

[0073] Step S20: determining at least one merging subtask according to the governance policy parameters, where the merging subtask corresponds to the partition table directory or the non-partition table directory to be merged.

[0074] Automatically determine the partition table directories or non-partition table directories that need to be merged based on the governance policy parameters, reducing the workload of manual identification and operation, reducing the risk of human error, and further improving the automation and efficiency of data management.

[0075] In this embodiment, the merging subtask may be to merge small files in the partition table directory or to merge small files in the non-partition table directory. The partition paths of the small files in the partition table directory and the small files in the non-partition table directory are different.

[0076] In an optional embodiment, the governance policy parameters include table name, partition importance mark and non-partition mark, and step S20 includes: determining the partition table directory to be merged based on the table name and partition importance mark; determining the non-partition table directory to be merged based on the table name and non-partition mark; generating at least one merge subtask based on the partition table directory or non-partition table directory to be merged.

[0077] It's important to note that by using table names and partition importance tags to identify partitioned table directories to merge, and by using table names and non-partition tags to identify non-partitioned table directories to merge, we can accurately locate the target directories to be merged, avoiding the processing of irrelevant directories and improving data management efficiency and accuracy. For partitioned tables, important partition data can be merged first based on importance tags to ensure the integrity and availability of critical data. For non-partitioned tables, targeted merge operations can also be performed to reduce data redundancy and improve data storage efficiency.

[0078] In an optional embodiment, the step of generating at least one merge subtask based on the partition table directory or non-partition table directory to be merged includes: obtaining directory information of the partition table directory or non-partition table directory to be merged, the directory information including the task identifier, table name and HDFS partition path; wherein the HDFS partition path includes the Hive root directory and the table partition path; and determining the merge subtask based on the directory information.

[0079] It should be noted that by obtaining detailed information about the directories to be merged, the specific content and scope of each merge task can be clearly identified. This helps to effectively manage and schedule numerous merge tasks in large-scale data processing scenarios, avoiding task confusion or omission.

[0080] Task IDs provide a unique identity for each merged subtask, making it easier to track and monitor tasks throughout the entire data processing process. This allows you to understand the execution status, progress, and results of each task in real time, allowing you to promptly identify and resolve issues that arise during task execution.

[0081] Obtaining HDFS partition paths enables the system to quickly and accurately locate the specific storage location of the data to be merged on HDFS. This includes the Hive root directory and table partition paths, enabling efficient navigation to the target data directory, reducing data lookup time and accelerating the startup of merge operations. Clear path information helps optimize data access strategies. For example, based on the path hierarchy and data distribution characteristics, the order and method of data reading can be rationally arranged, reducing network overhead and disk I / O (Input / Output) operations during data transmission, thereby improving data access efficiency.

[0082] Optionally, a merged subtask set is constructed based on at least one merged subtask. Figure 4 After data preparation is complete, a DataFrame task set is constructed based on the data structure of the task ID, table name, and HDFS partition path. The HDFS partition path includes the Hive root directory and the table partition path. In the data display section, the path structure of the partitioned tables tableA and tableB is slightly different from that of the non-partitioned table tableM.

[0083] Step S30: Determine the number of Spark computing tasks corresponding to at least one of the merged subtasks.

[0084] In this embodiment, the Spark computing task is a serial execution process that merges subtasks, and the Spark computing tasks do not interfere with each other.

[0085] Optionally, the number of Spark computing tasks is determined based on the number of merged subtasks, and the merged subtasks corresponding to the Spark computing tasks are allocated based on the number of Spark computing tasks. For example, the number of merged subtasks is 30, the number of Spark computing tasks is 10, and each Spark computing task is allocated 3 merged subtasks.

[0086] Optionally, the number of Spark computing tasks is determined according to a preset concurrency configuration threshold, and the merged subtasks are grouped according to the number of Spark computing tasks.

[0087] Step S40: grouping the merged subtasks according to the number of the Spark computing tasks, and constructing a corresponding Spark computing task set according to the grouped merged subtasks, wherein each Spark computing task is used to serially process the assigned merged subtasks.

[0088] Leveraging Spark's distributed computing framework, merging subtasks is assigned to multiple Spark computing tasks for parallel processing, significantly reducing data processing time and improving data processing performance and throughput, enabling faster fulfillment of business requirements for timely data processing. By grouping merging subtasks and constructing Spark computing task collections, task scheduling and resource allocation can be more flexibly performed based on task characteristics and resource availability, further optimizing the data processing process and performance.

[0089] In this embodiment, after a Spark computing task is constructed, the merge subtask is distributed to the Spark computing task. Each Spark computing task corresponds to at least one merge subtask. For example, one Spark computing task corresponds to four merge subtasks. Merge tasks between different Spark computing tasks do not interfere with each other and are performed independently. The Spark computing task is used to merge small files in the partition table directory or non-partition table directory corresponding to the assigned merge subtask.

[0090] In one embodiment, referring to Figure 2The file processing solution is generally divided into the following four steps: Step 1: Task Preparation. Redundant partitions are preprocessed, valid partitions are retained, and valid partitioned and non-partitioned tables are merged into the same set. The merge concurrency is calculated based on the set and concurrency parameter configuration. An accumulator is initialized, which is responsible for counting merge progress and results to facilitate post-merge partition metadata repair. Step 2: Task Set Construction. A Dataset task dataset is constructed based on the merge list in Step 1. Each row represents a merge subtask, corresponding to a file in a partitioned table directory or a file in a non-partitioned table directory. Progress monitoring threads and task threads are simultaneously started. A row is the smallest data unit in Spark data structures, representing a row of structured data. Step 3: Task Distribution and Processing. The subtasks in Step 2 are distributed to various Spark tasks on the executor side. Multiple Spark tasks are executed concurrently, with the number of tasks equal to the concurrency configured in Step 1. Each Spark task contains one or more merge subtasks. These merge subtasks call the MergeFiles method in OrcFiles to operate on the Hive partitioned table directory in HDFS to merge small files. Step 4: Task Feedback. After the merge subtask completes, the accumulator is updated. A listening thread in the driver monitors the accumulator status in real time to obtain real-time merge progress. After all tasks are completed, a return is made to the driver, which repairs HDFS metadata based on the accumulator status and prints the merge results. Throughout this process, the legality verification module ensures data compliance at every stage.

[0091] In the technical solution of this embodiment, the number of corresponding Spark computing tasks is determined based on the merged subtasks, and the merged subtasks are grouped to construct a corresponding Spark computing task set, thereby achieving reasonable allocation and utilization of computing resources. Each Spark computing task can focus on serial processing of the merged subtasks assigned to it, avoiding excessive waste or insufficiency of resources and improving the utilization of cluster resources. For duplicate or redundant data in partitioned tables and non-partitioned tables, the processing of merged subtasks can be effectively integrated and cleaned up, reducing data storage redundancy and storage costs. At the same time, it also avoids repeated calculation of duplicate data during data processing and improves management efficiency.

[0092] Based on any of the above embodiments of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be described in detail later. Figure 5 , step S30 includes:

[0093] Step S31, determining a first configuration concurrency number according to the concurrency configuration threshold and the total number of the merged subtasks;

[0094] Step S32: Determine the number of Spark computing tasks according to the first configured concurrency.

[0095] Setting the concurrency rate too high can lead to excessive consumption of system resources, such as memory, causing system crashes or task execution failures. By determining an appropriate concurrency rate based on the concurrency configuration threshold, this situation can be effectively avoided, ensuring stable system operation. Properly determining the concurrency rate helps balance the load across nodes in the cluster, preventing some nodes from being overloaded while others are underloaded due to an unreasonable concurrency rate, thereby improving cluster reliability and stability.

[0096] Optionally, a first configured concurrency number is determined based on the minimum value of the concurrency configuration threshold and the total number of merged subtasks, and the first configured concurrency number is used as the number of Spark computing tasks.

[0097] In another optional embodiment, the second configured concurrency is determined based on the number of available cores in the cluster in the computing resources; the minimum configured concurrency between the first configured concurrency and the second configured concurrency is determined; and the number of Spark computing tasks is determined based on the minimum configured concurrency.

[0098] It should be noted that the concurrency configuration threshold reflects the concurrency requirements of the business logic and data processing complexity. Determining the first configured concurrency based on the total number of merged subtasks and this threshold ensures that allocated resources meet task processing requirements while avoiding resource waste caused by overallocation. The number of available cores in the cluster represents the physical computing resource constraints. By comprehensively considering these two factors, the first and second configured concurrency values ​​are determined to be more comprehensive and accurate, better matching the actual available computing resources and avoiding the situation where the concurrency setting is too high or too low based on only one factor. Taking the minimum of the first and second configured concurrency values ​​as the basis for the final determination of the number of Spark computing tasks ensures that the data processing logic requirements are met while not exceeding the actual concurrency capacity supported by the cluster. This minimizes resource waste and improves resource utilization efficiency, while also preventing performance issues caused by excessive cluster resource consumption due to overly high concurrency settings.

[0099] In the technical solution of this embodiment, the concurrency configuration threshold provides a clear reference standard for resource allocation. The first configuration concurrency is determined based on the total number of merged subtasks and the threshold, which can ensure that the allocated resources meet the task processing requirements while avoiding waste of resources caused by excessive allocation. For example, if the cluster resources are limited and there are many merged subtasks, this method can accurately allocate appropriate concurrency to the tasks instead of blindly using too many resources. By determining the number of Spark computing tasks based on the first configuration concurrency, each computing task can use the allocated resources more efficiently. Compared with arbitrarily determining the number of computing tasks, this method can ensure that resources are fully utilized and improve the overall operating efficiency of the cluster.

[0100] Based on any of the above embodiments of the present application, in the third embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be described in detail later. Figure 6 , after step S40, further comprising:

[0101] Step S50: Distribute the Spark computing task set to the executor, so that the executor calls the preset algorithm of the ORC engine to merge the target files in the partition table directory or the non-partition table directory for each of the merge subtasks, and the target file size is less than a preset threshold.

[0102] In this embodiment, for distributed computing frameworks such as Spark, a large number of tasks will be generated according to the number of files or partitions during task scheduling. Too many small files will lead to an excessive number of tasks, increasing the complexity and overhead of task scheduling. By merging files to reduce the number of tasks, the delay in task scheduling can be reduced and the overall resource utilization efficiency can be improved. Target files, that is, small files, in big data systems specifically refer to data files that are much smaller than the block size of the distributed file system, and their existence will significantly reduce system performance. The size of small files is much smaller than the HDFS block configuration, usually ≤128MB, the default block size. The size of small files is lower than the business preset threshold, such as the preset threshold dynamically set through the Hive governance configuration table. The scenarios for the generation of small files can be Hive tables with high frequency writes, streaming job outputs, and over-partitioned Spark results.

[0103] In one embodiment, referring to Figure 7Based on the constructed Dataset task dataset and the configured concurrency, the Driver distributes merge tasks to Spark Tasks on the Executor side. The number of Spark Tasks is equal to the configured concurrency, and each Spark Task corresponds to one or more merge subtasks, Rows. Each Spark Task sequentially processes all internal merge subtasks, ensuring that multiple Spark Tasks do not interfere with each other and concurrently process different sets of subtasks. Alternatively, the mergeFiles method of the OrcFile class in the org.apache.orc package can be used to merge small files. This method can merge multiple Orc files with the same directory schema into a single Orc file and also merge Hive user metadata, ensuring good compatibility with Orc-formatted tables. Small file merging is implemented on HDFS through the Spark Executor side. This allows the Dataset itself to contain multiple merge tasks, which are then dispatched to the Executor side for high-concurrency processing based on the Spark Task dimension. The configured concurrency of the Spark Task is flexibly configurable. When partitioning a large number of small file tables is required, the configured concurrency can be appropriately increased.

[0104] Alternatively, the mergeFiles method of the OrcFile class in the org.apache.orc package can be used to merge small files. This allows multiple Orc files with the same directory schema to be merged into a single Orc file, while also incorporating Hive usermetadata, ensuring good compatibility with Orc-formatted tables. Small file merging can be implemented on HDFS through the Spark Executor. This allows the Dataset itself to contain multiple merge tasks, which are then dispatched to the Executor for highly concurrent processing as Spark Tasks. The Spark Task concurrency can be flexibly configured; when partitioning a large number of small file tables is required, the concurrency can be appropriately increased.

[0105] In an optional embodiment, during the merge process, a temporary directory is created under the target directory, and the temporary directory is used to store the merge results; after the merge is successful, the temporary directory replaces the partition table directory or the non-partition table directory. During the merge process, direct operations on the partition table directory or the non-partition table directory may cause data read and write conflicts. For example, when other processes are reading or writing files in these directories, the merge operation may cause data inconsistency or read and write errors. By creating a temporary directory to store the merge results, direct modification of the original directory can be avoided, and the replacement is performed after the merge is completed to ensure the consistency and integrity of data reading and writing. Performing the merge operation directly in the original directory may lead to frequent file creation, deletion, and modification operations, thereby increasing the number of changes to the file system metadata. The use of a temporary directory can reduce the metadata operations on the original directory. After the merge is completed, the directory structure is updated at one time through the replacement operation, which optimizes the operating efficiency of the file system and reduces the overhead of metadata management.

[0106] The merge creates a new temporary directory, tmp, under the Hive partition directory on HDFS. If the merge fails for small files, it only affects the newly merged small files in the tmp directory. It does not affect the actual read / write directory, nor does it affect the normal reading of table data during the merge. If the merge succeeds, each file operation strategy ensures that the original data is deleted after the operation is completed. During deletion, the validity verification service is used to perform strong verification on the deleted objects to ensure data security.

[0107] In an optional embodiment, the merge progress and merge results of each Spark computing task are monitored in real time based on an accumulator; the partitioned table or non-partitioned table that has been successfully merged is determined based on the merge progress and merge results; and a metadata repair operation is performed on the partitioned table or non-partitioned table that has been successfully merged.

[0108] By using accumulators to monitor the merge progress of each Spark task in real time, you can keep up to date on the execution status of each task, including the amount of merged work completed and the amount remaining. This helps you dynamically monitor the entire merge process, identifying potential delays or anomalies so you can take timely adjustments or optimization measures. Real-time monitoring of merge results quickly reveals whether each Spark task has successfully completed the merge operation and whether the merged data meets expectations. Failed merge tasks can be promptly investigated and re-executed to prevent the accumulation and spread of errors and ensure the accuracy and completeness of the merge.

[0109] In one embodiment, referring to Figure 8When a single merge subtask completes, the executor updates the merge table and partitions to the corresponding Long-valued accumulator and custom List accumulator. The current progress accumulator increments by 1 for all merges. When all tasks are processed, the driver automatically repairs the HDFS directory metadata for the successfully merged partitions.

[0110] In the technical solution of this embodiment, by merging small files in the partition table directory or the non-partition table directory, the number of files can be significantly reduced and the file reading and writing efficiency can be improved. Too many files will lead to a large number of random read and write operations, increase disk I / O overhead, and affect overall performance. The merged files can make full use of the advantages of sequential reading and writing of files, reduce disk I / O latency, and improve query and write performance. The preset algorithm of the ORC engine can efficiently merge files, and the ORC file format itself has good compression characteristics and columnar storage structure. For query operations, especially column filtering and aggregation queries, the ORC format can quickly locate and read the required data columns, reduce the amount of data scanning, and improve the execution speed of queries.

[0111] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the document processing method of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0112] The present application provides a file processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the file processing method in the above-mentioned embodiment 1.

[0113] Reference below Figure 9 , which shows a schematic structural diagram of a file processing device suitable for implementing an embodiment of the present application. The file processing device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The file processing device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0114] like Figure 9As shown, the file processing device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the file processing device are also stored in RAM 1004. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. Communication device 1009 can allow the file processing device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a file processing device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0115] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0116] The file processing device provided by this application, employing the file processing method of the aforementioned embodiment, can solve the technical problem of coordinating and managing small files. Compared to the prior art, the beneficial effects of the file processing device provided by this application are the same as those of the file processing method provided by the aforementioned embodiment. Other technical features of the file processing device are the same as those disclosed in the aforementioned embodiment and are not further described here.

[0117] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0118] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0119] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, wherein the computer-readable program instructions are used to execute the file processing method in the above embodiment.

[0120] The computer-readable storage medium provided in this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF, Radio Frequency), etc., or any suitable combination thereof.

[0121] The computer-readable storage medium may be included in the file processing device, or may exist independently without being assembled into the file processing device.

[0122] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the file processing device, the file processing device determines the number of corresponding Spark computing tasks based on the merged subtasks, and groups the merged subtasks to construct a corresponding Spark computing task set, thereby achieving reasonable allocation and utilization of computing resources. Each Spark computing task can focus on serial processing of the assigned merged subtasks, avoiding excessive waste or insufficiency of resources and improving the utilization of cluster resources. For duplicate or redundant data in partitioned tables and non-partitioned tables, the merged subtasks can be effectively integrated and cleaned up, reducing data storage redundancy and storage costs. At the same time, it also avoids repeated calculation of duplicate data during data processing and improves management efficiency.

[0123] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0124] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0126] The computer-readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned file processing method, thereby solving the technical problem of comprehensively managing small files. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the file processing method provided in the aforementioned embodiment, and are not further elaborated here.

[0127] The present application also provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned file processing method when executed by a processor.

[0128] The computer program product provided in this application can solve the technical problem of coordinating and managing small files. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the file processing method provided in the above embodiment, and will not be repeated here.

[0129] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A file processing method, characterized in that: The file processing method includes: Constructing a Hive governance configuration table, which includes governance policy parameters for partitioned tables and non-partitioned tables; Determine at least one merging subtask according to the governance policy parameters, where the merging subtask corresponds to the partition table directory or the non-partition table directory to be merged; Determine the number of Spark computing tasks corresponding to at least one of the merged subtasks; The merged subtasks are grouped according to the number of the Spark computing tasks, and a corresponding Spark computing task set is constructed according to the grouped merged subtasks. Each Spark computing task is used to serially process the assigned merged subtasks.

2. The file processing method according to claim 1, wherein: The governance policy parameters include a concurrency configuration threshold, and the step of determining the number of Spark computing tasks corresponding to at least one of the merged subtasks includes: Determining a first configuration concurrency number according to the concurrency configuration threshold and the total number of the merged subtasks; The number of Spark computing tasks is determined according to the first configured concurrency number.

3. The file processing method according to claim 2, wherein: The step of determining the number of Spark computing tasks according to the first configured concurrency number includes: Determine the second configuration concurrency based on the number of available cores in the cluster; Determine the minimum configured concurrency number between the first configured concurrency number and the second configured concurrency number; The number of Spark computing tasks is determined based on the minimum configured concurrency.

4. The file processing method according to claim 1, wherein: The governance policy parameters include a table name, a partition importance tag, and a non-partition tag. The step of determining at least one merging subtask according to the governance policy parameters includes: Determine the partition table directory to be merged according to the table name and the partition importance mark; Determine a non-partition table directory to be merged according to the table name and the non-partition mark; At least one merging subtask is generated according to the partition table directory to be merged or the non-partition table directory.

5. The file processing method according to claim 4, wherein: The step of generating at least one merging subtask according to the partition table directory to be merged or the non-partition table directory includes: Obtain directory information of the partition table directory or the non-partition table directory to be merged, wherein the directory information includes a task identifier, a table name, and an HDFS partition path; wherein the HDFS partition path includes a Hive root directory and a table partition path; The merging subtask is determined according to the directory information.

6. The file processing method according to claim 1, wherein: After the step of constructing a corresponding Spark computing task set according to the grouped merged subtasks, the method further includes: The Spark computing task set is distributed to an executor, so that the executor calls a preset algorithm of the ORC engine to merge the target files in the partition table directory or the non-partition table directory for each of the merge subtasks, and the size of the target file is smaller than a preset threshold.

7. The file processing method according to claim 6, wherein: After the step of distributing the Spark computing task set to the executor, the method further includes: During the merging process, a temporary directory is created under the target directory, and the temporary directory is used to store the merging results; After the merge is successful, the temporary directory replaces the partition table directory or the non-partition table directory.

8. The file processing method according to claim 1, wherein: After the step of distributing the Spark computing task set to the executor, the method further includes: Based on the accumulator, real-time monitoring is performed on the merging progress and merging results of each Spark computing task; Determining the partition table or the non-partition table that has been successfully merged according to the merge progress and the merge result; A metadata repair operation is performed on the successfully merged partition table or the non-partition table.

9. A file processing device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the file processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the file processing method according to any one of claims 1 to 8.