File migration method, storage medium and electronic device
By creating a target table in a Hive table and automatically setting parameters using the target mapping relationship, the problem of low file migration efficiency caused by different Hive table permissions and sizes is solved, and efficient file migration and format conversion are achieved.
Patent Information
- Application Number
- CN202510811237.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, due to different permissions and sizes of Hive tables, MapReduce tasks that perform compression format conversion operations require manual parameter setting, resulting in high costs, slow speed and high failure rate.
By creating a target table and determining the target conversion parameter values in the target mapping relationship based on the basic information of the source table and the data volume of the files to be migrated, the parameters are automatically set for file format conversion and migration by using the correspondence between the file data volume and the conversion parameter values within the historical time range recorded in the target mapping relationship.
It realizes the automated file migration process, improves the efficiency and reliability of file migration, and reduces resource waste and business interruption.
Smart Images

Figure CN120743875A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a method for migrating files, a storage medium, and an electronic device. Background Art
[0002] As big data becomes increasingly widely used across industries, the volume and variety of data are increasing, and the demand for data processing is also growing. Hive tables are a data warehouse tool within Hadoop that primarily maps structured data files into database tables, providing users with simple SQL query capabilities. Using Hive tables for storing, querying, and analyzing large data sets can greatly improve the convenience and efficiency of big data processing.
[0003] In practical applications, Hive tables can support structured data files with different compression formats. Conversion of the compression format of data files in Hive tables typically involves creating new tables with the same fields but different compression formats. These tables are then written to the new tables through query rewriting to achieve the compression format conversion. However, due to the varying permissions and sizes of different Hive tables, the MapReduce tasks performing the compression format conversion require different parameter values. Existing parameter setting methods often rely on manual experience, sequentially setting parameters for MapReduce tasks for different Hive tables. This results in high compression format conversion costs, slow conversion speeds, and a high conversion failure rate for data files in Hive tables.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] Embodiments of the present invention provide a method for migrating files, a storage medium, and an electronic device to at least solve the problem in the related art that files cannot be automatically migrated due to different parameter settings required for file migration in various tables.
[0006] According to one embodiment of the present invention, a method for migrating files is provided, comprising: creating a target table based on basic information of a source table, wherein the source table and the target table support different file formats for storage; when migrating files in the source table to the target table, determining a target conversion parameter value in a target mapping relationship based on the data volume of the files to be migrated, wherein the target mapping relationship records the correspondence between the data volume of the files migrated within a historical time range and the conversion parameter value; performing format conversion on the files to be migrated using the target conversion parameter value, and migrating the files to be migrated after the format conversion to the target table.
[0007] In an exemplary embodiment, a target conversion parameter value is determined in a target mapping relationship based on the data volume of the file to be migrated, including: determining a target data volume interval corresponding to the data volume of the file to be migrated in the target mapping relationship, wherein the target mapping relationship includes at least one data volume interval, each of the data volume intervals corresponds to multiple conversion migration durations, and each conversion migration duration corresponds to one conversion parameter value; determining a target conversion migration duration from the multiple conversion migration durations corresponding to the target data volume interval; and determining the conversion parameter value corresponding to the target conversion migration duration as the target conversion parameter value.
[0008] In an exemplary embodiment, a target conversion migration duration is determined from among the multiple conversion migration durations corresponding to the target data volume interval, including: determining the smallest conversion migration duration among the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, determining the conversion migration duration that is in the middle of the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, in the case where the target data volume interval corresponds to N conversion migration durations, sorting the N conversion migration durations to obtain the sorted N conversion migration durations; dividing the target data volume interval into N sub-data volume intervals, and establishing a correspondence between each sub-data volume interval and each sorted conversion migration duration in the order of the sorted N conversion migration durations; determining the sub-data volume interval in which the data volume of the file to be migrated is located as the target sub-data volume interval, and determining the conversion migration duration corresponding to the target sub-data volume interval as the target conversion migration duration.
[0009] In an exemplary embodiment, migrating the format-converted files to be migrated to the target table includes: determining the data volume of files in each partition of the source table, wherein the source table includes multiple partitions; determining a partition in the source table having a non-zero data volume of files as a first target partition; and migrating the files to be migrated in the first target partition to the target table.
[0010] In an exemplary embodiment, migrating the files to be migrated in the first target partition to the target table includes: creating a second target partition in the target table; migrating the files to be migrated in the first target partition to the second target partition, wherein the paths of the first target partition and the second target partition are the same.
[0011] In an exemplary embodiment, the method further includes: determining a partition in the source table whose file data volume is zero as a third target partition; and creating a fourth target partition in the target table, wherein the fourth target partition is an empty partition and has the same path as the third target partition.
[0012] In an exemplary embodiment, after migrating the format-converted file to be migrated to the target table, the method further includes: determining the duration of the format conversion of the file to be migrated as the current conversion duration; determining the duration of migrating the format-converted file to be migrated to the target table as the current migration duration; determining the sum of the current conversion duration and the current migration duration as the current conversion migration duration; and updating the target mapping relationship according to the target conversion parameter value and the current conversion migration duration.
[0013] In an exemplary embodiment, the target mapping relationship is updated according to the target conversion parameter value and the current conversion migration duration, including: when the current conversion migration duration is included in the multiple conversion migration durations corresponding to the target data volume interval, the conversion parameter value corresponding to the current conversion migration duration in the target mapping relationship is updated to the target conversion parameter value; when the current conversion migration duration is not included in the multiple conversion migration durations corresponding to the target data volume interval, an association relationship between the current conversion migration duration, the target conversion parameter value and the target data volume interval is established in the target mapping relationship.
[0014] According to another embodiment of the present invention, a device for migrating files is provided, comprising: a creation module for creating a target table based on basic information of a source table, wherein the source table and the target table support different file formats for storage; a determination module for determining a target conversion parameter value in a target mapping relationship based on the data volume of the file to be migrated when migrating the file in the source table to the target table, wherein the target mapping relationship records the correspondence between the data volume of the file migrated within a historical time range and the conversion parameter value; a migration module for performing format conversion on the file to be migrated using the target conversion parameter value, and migrating the file to be migrated after the format conversion to the target table.
[0015] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0016] According to another embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.
[0017] According to yet another embodiment of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0018] Through the present invention, a target table with a different file format from that stored in the source table is created based on basic information of a source table. When migrating files to be migrated in the source table to the target table, a target conversion parameter value corresponding to the data volume is determined in a target mapping relationship based on the data volume of the files to be migrated in the source table. Thus, the files to be migrated in the source table are format-converted according to the target conversion parameter value, and the files to be migrated after the format conversion are migrated from the source table to the target table.
[0019] Because the target mapping relationship records the correspondence between the data volume of migrated files within a historical time range and the conversion parameter values, when performing format conversion and migration on the files to be migrated in the source table, the optimal conversion parameter values for the current format conversion and migration task can be quickly determined in the target mapping relationship based on the data volume of the files to be migrated in the source table. Parameters can then be set based on the determined optimal conversion parameter values to perform format conversion and migration on the files to be migrated in the source table. This solves the problem in the related art of being unable to automatically migrate files due to different parameter settings required for file migration in each table, thereby improving the efficiency of file migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a hardware structure block diagram of a mobile terminal according to a method for migrating files according to an embodiment of the present invention;
[0021] Figure 2 is a flowchart of a method for migrating files according to an embodiment of the present invention;
[0022] Figure 3 This is a flowchart of batch conversion of Hive tables into compression formats according to an embodiment of the present invention;
[0023] Figure 4 It is a structural block diagram of an apparatus for migrating files according to an embodiment of the invention. DETAILED DESCRIPTION
[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.
[0025] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0026] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG is a hardware structure block diagram of a mobile terminal of a method for migrating files according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0027] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for migrating files in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0028] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by the mobile terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0029] In this embodiment, a method for migrating files running on the above mobile terminal or network architecture is provided. Figure 2 is a flow chart of a method for migrating files according to an embodiment of the present invention. Figure 2As shown, the process includes the following steps:
[0030] Step S202: creating a target table based on basic information of the source table, wherein the source table and the target table support different file formats for storage;
[0031] The above-mentioned source table can be a table mapped to the structured data files stored in the data warehouse, such as Hive table, relational database table, BigQuery table, and Teradata table. By mapping the structured data files stored in the database into the table structure, the structured data files can be managed and utilized more efficiently, thereby supporting more complex data analysis and business decision-making.
[0032] The above basic information may be the table name, field (or column), data file, partition, storage path of the corresponding data in the data warehouse, etc. of the source table. Among them, the data file in the source table may be a data file compressed by a compression algorithm. The specific file format of the compressed data file may be snappy compression format, gzip compression format, bzip2 compression format, zlib compression format, etc.
[0033] The target table can have the same fields and partitions as the source table, but a different data file format. For example, if the source table's data files are in snappy compression format, the target table's data files can be in a compression format other than snappy, such as zlib.
[0034] Taking a Hive table as an example, we obtain basic information about the original Hive table (hereinafter referred to as the original table, also known as the source table), such as the original table's fields, attributes (including file format and compression format, etc.), partitions, the original table's corresponding HDFS (Hadoop Distributed File System) data path, and each partition path. Based on this basic information, we create a new Hive table (hereinafter referred to as the new table, also known as the target table) with the same fields and partition settings as the original table but a different compression format.
[0035] Step S204: When migrating files from the source table to the target table, determine target conversion parameter values in a target mapping relationship based on the data volume of the files to be migrated, wherein the target mapping relationship records the correspondence between the data volume of the files migrated within a historical time range and the conversion parameter values;
[0036] The files to be migrated may be data files in the source table, and the file format of the data files is a first compression format, such as snappy compression format, gzip compression format, bzip2 compression format, zlib compression format, etc.
[0037] The target mapping relationship may be a mapping relationship between the file data volume, the conversion migration duration of the file of that data volume, and the conversion parameter value. The mapping relationship may be determined based on the actual situation of the file migration tasks performed within a historical time range.
[0038] Among them, the data volume of the file can be the specific data volume of the converted and migrated data file, that is, the specific data volume of the data file contained in the original table, such as 10T, 50T, 100T, etc.; the file conversion and migration time can be the total time required to convert, compress, format and migrate the data file in the original table to the new table. The conversion and migration time corresponding to data files with different data volumes is different; the conversion parameter value can be the parameter of the model or algorithm used to control the above operations when performing the compression format conversion operation of the data file and the migration operation of the compressed format data file. Taking the MapReduce task parameter as an example, the specific conversion parameter value can include the number of Map tasks (mapreduce.job.maps) and the number of Reduce tasks (mapreduce.job.reduces) of the MapReduce task, which are used to determine the parallelism of data processing. Too many or too few Map or Reduce tasks may affect the processing efficiency and the Application when executing the MapReduce task. The Master resource size (am.resource.mb) and Reducer memory size (reduce.memory.mb) control the computing resources available for each task and the file split size (split.minsize / maxsize) of MapReduce tasks, ensuring the parallel processing capability and data reading efficiency of tasks. Optimizing the conversion parameter values can help improve task execution efficiency, reduce resource waste, and avoid task failures due to insufficient resources. Proper adjustment of these parameters is particularly necessary when processing large data sets and performing complex computing tasks.
[0039] The target conversion parameter value may be the parameter of the model or algorithm that needs to be set for the current file migration task, that is, the compression format conversion operation of the data file and the migration operation of the compressed format-converted data file.
[0040] Taking the Hive table as an example, when it is determined that the data files in the original table need to be converted to a compressed format and migrated to a new table, the data volume of the data files in the original table (i.e., the files to be migrated) is obtained, and the conversion parameter values corresponding to the data files of this data volume within the historical time range are determined in the target mapping relationship. These values are used as the conversion format of the current data files and the target conversion parameter values for the migration task.
[0041] Optionally, a target conversion parameter value is determined in a target mapping relationship based on the data volume of the file to be migrated, including: determining a target data volume interval corresponding to the data volume of the file to be migrated in the target mapping relationship, wherein the target mapping relationship includes at least one data volume interval, each of the data volume intervals corresponds to multiple conversion migration durations, and each conversion migration duration corresponds to one conversion parameter value; determining a target conversion migration duration from the multiple conversion migration durations corresponding to the target data volume interval; and determining the conversion parameter value corresponding to the target conversion migration duration as the target conversion parameter value.
[0042] Table 1 below is an example of a target mapping relationship using a Hive table as an example. As shown in Table 1, the target mapping relationship contains at least one data volume interval, each data volume interval corresponds to multiple conversion and migration durations, and each conversion and migration duration corresponds to a set of conversion parameter values. The data volume of the file to be migrated in the source table is determined, and the target data volume interval where the data volume of the file to be migrated is located is found in the target mapping relationship. Then, the target conversion parameter values required for optimizing the compression format conversion and migration operations for the migrated file are determined from the multiple conversion and migration durations corresponding to the target data volume interval, and the parameters of the model or algorithm for performing the compression format and migration operations are set according to the target conversion parameter values. Through the above content, selecting appropriate conversion and migration durations and parameter values according to the specific data volume can better optimize the performance of compression and migration, reduce unnecessary calculations, and improve overall operational efficiency. At the same time, through the target mapping relationship, it is possible to adapt to different data environments and requirements, automatically select the optimal conversion parameter values, flexibly respond to various complex data migration scenarios, and effectively improve the efficiency and reliability of file migration.
[0043] Table 1 Target mapping relationship
[0044]
[0045] Step S206 : performing format conversion on the file to be migrated according to the target conversion parameter value, and migrating the file to be migrated after the format conversion to the target table.
[0046] After determining the target conversion parameter value, the parameters of the algorithm or model that performs the file format conversion task and the migration task are set according to the target conversion parameter value, so that the files to be migrated in the source table are converted into formats and then migrated to the target table through the model or algorithm whose parameters are set to the target conversion parameter value.
[0047] Specifically, migrating the to-be-migrated files after format conversion to the target table includes: determining the data volume of the files in each partition of the source table, wherein the source table includes multiple partitions; determining a partition in the source table having a non-zero data volume of the files as a first target partition; and migrating the to-be-migrated files in the first target partition to the target table.
[0048] Partitioning can be used to group data in a raw table by the values of one or more columns. Different groups (or partitions) are stored in different directories. For example, in a Hive table, partitions can group data by specific attributes, and different partitions have different physical storage locations on HDFS. Table partitioning is a key step in efficiently reading and rewriting data, especially during data compression format conversion. It supports efficient processing of massive amounts of data, improving data query efficiency and management capabilities.
[0049] The first target partition can be a partition in the source table containing non-zero file data. When migrating the converted files, the partition structure of the source table is first determined. The partition containing non-zero file data in the source table is then designated as the first target partition. The files to be migrated from the first target partition are then migrated to the corresponding partition in the target table. Partitioning the original table ensures that the converted new table inherits the structure of the original table and that the physical location of the data is correctly updated, maintaining data consistency and integrity.
[0050] Optionally, the execution entity of the above steps can be a background processor, or other devices with similar processing capabilities, or a machine that integrates at least an image acquisition device and a data processing device, wherein the image acquisition device may include a graphics acquisition module such as a camera, and the data processing device may include a computer, a mobile phone and other terminals, but is not limited to this.
[0051] Through the above steps, a target table is created based on the basic information of the source table, having the same attributes and partition structure as the source table but storing a different file format. When migrating the files to be migrated in the source table to the target table, the target conversion parameter value corresponding to the data volume of the files to be migrated in the source table is determined in the target mapping relationship, thereby setting the parameters of the model or algorithm for executing the file format conversion and migration tasks based on the target conversion parameter value. After the parameter setting, the model or algorithm performs format conversion on the files to be migrated in the source table, and then migrates the format-converted files to be migrated from the source table to the target table. Because the target mapping relationship records the correspondence between the data volume of the files migrated within a historical time range and the conversion parameter value, when format conversion and migration are performed on the files to be migrated in the source table, the optimal conversion parameter value for the current format conversion and migration task can be quickly determined in the target mapping relationship based on the data volume of the files to be migrated in the source table. Parameters are then set based on the determined optimal conversion parameter value, and the files to be migrated in the source table are format converted and migrated. This solves the problem in the related art that file migration cannot be automatically migrated due to different parameter settings required for file migration in each table, thereby improving the efficiency of file migration.
[0052] As an optional implementation, determining a target conversion migration duration from among the multiple conversion migration durations corresponding to the target data volume interval includes: determining the smallest conversion migration duration from among the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, determining the conversion migration duration that is in the middle of the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, in the case where the target data volume interval corresponds to N conversion migration durations, sorting the N conversion migration durations to obtain the sorted N conversion migration durations; dividing the target data volume interval into N sub-data volume intervals, and establishing a correspondence between each sub-data volume interval and each sorted conversion migration duration in the order of the sorted N conversion migration durations; determining the sub-data volume interval in which the data volume of the file to be migrated is located as the target sub-data volume interval, and determining the conversion migration duration corresponding to the target sub-data volume interval as the target conversion migration duration.
[0053] In the target mapping relationship, each data volume interval corresponds to multiple conversion migration durations, and each conversion migration duration corresponds to a conversion parameter value. Therefore, when determining the target conversion parameter value, it is necessary to determine the most appropriate target conversion relationship value from the multiple conversion parameter values corresponding to the target data volume interval. Alternatively, the most appropriate target conversion relationship value can be determined from the multiple conversion parameter values corresponding to the target data volume interval based on the conversion relationship duration.
[0054] For example, the smallest conversion migration duration among the multiple conversion migration durations corresponding to the target data volume interval is determined as the target conversion migration duration, and then the conversion parameter value corresponding to the target conversion migration duration is determined as the target conversion parameter value, so as to achieve fast and efficient migration, minimize the migration time, improve user experience, and reduce business interruptions caused by long migration.
[0055] Alternatively, the conversion migration duration with the median ranking among the multiple conversion migration durations corresponding to the target data volume interval is determined as the target conversion migration duration, and then the conversion parameter value corresponding to the target conversion migration duration is determined as the target conversion parameter value. By sorting multiple conversion migration durations and selecting the median, the influence of extreme values on the results can be effectively reduced, thereby obtaining a more robust and reliable target conversion migration duration, ensuring that in most cases, the conversion migration process is efficient and meets expectations.
[0056] Alternatively, the target data volume interval is divided twice according to the data volume of the conversion migration duration corresponding to the target data volume interval to obtain multiple sub-data volume intervals, and then the target conversion migration duration is determined from the multiple conversion migration durations corresponding to the target data volume interval according to the target sub-data volume interval where the data volume of the file to be migrated is located, and then the conversion parameter value corresponding to the target conversion migration duration is determined as the target conversion parameter value. For example, if the target data volume interval corresponds to 5 conversion migration durations, the target data volume interval (1-10T) is divided into 5 parts to obtain 5 sub-data volume intervals of 1-2T, 3-4T, 5-6T, 7-8T, and 9-10T respectively. At the same time, the 5 conversion migration durations are sorted as 1.5h, 1.8h, 2h, 2.5h, and 3h. Therefore, each sub-data volume interval corresponds to a conversion migration duration, such as 1-2T corresponds to 1.5h, 3-4T corresponds to 1.8h, 5-6T corresponds to 2h, etc. It is determined that the data volume of the file to be migrated is 5T and is located in the third sub-data volume interval 5-6T. Then the conversion migration duration of 2h corresponding to the target data volume interval 5-6T is determined as the target conversion migration duration, and then the conversion parameter value corresponding to the target conversion migration duration is determined as the target conversion parameter value. Through precise division and matching of data volume, the migration duration closest to the actual demand can be selected. More precise control helps reduce the risk of errors or failures during the migration process. Adjustment and expansion can be made according to different system requirements and environmental changes, thereby optimizing the migration process and reducing unnecessary waiting time and resource waste.
[0057] As an optional implementation, migrating the files to be migrated in the first target partition to the target table includes: creating a second target partition in the target table; migrating the files to be migrated in the first target partition to the second target partition, wherein the paths of the first target partition and the second target partition are the same.
[0058] During the process of migrating the files to be migrated after format conversion in the first target partition to the target table, the position of the first target partition in the source table is determined, and a second target partition is created corresponding to the same position in the target table. The second target partition has the same storage file path as the first target partition, thereby migrating the files to be migrated after format conversion in the first partition to the second target partition, thereby achieving the migration of files in non-empty partitions (i.e., partitions with non-zero data volume) in the original table.
[0059] As an optional implementation, the method further includes: determining a partition in the source table whose file data volume is zero as a third target partition; creating a fourth target partition in the target table, wherein the fourth target partition is an empty partition, and the path of the fourth target partition is the same as that of the third target partition.
[0060] The third target partition can be a partition with zero data volume in the source table. For an empty partition in the source table (i.e., a partition with zero data volume), it is only necessary to determine the position of the third target partition in the source table and create a fourth target partition at the same position in the target table. The storage file path of the fourth target partition is the same as that of the third target partition, and the fourth target partition is also an empty partition, ensuring the consistency of the original table and the target table, avoiding the unavailability of the entire table due to partition damage, and thus improving data availability.
[0061] As an optional implementation, before converting the format of the file to be migrated using the target conversion parameter value and migrating the converted file to the target table, the method further includes: obtaining permission information of the source table; and setting the target table according to the permission information.
[0062] By setting the permissions of the target table to be consistent with the permissions of the source table, it is helpful to achieve automation and consistency in database permission management. Especially in large systems, it is an important guarantee for ensuring the security and stable operation of the system.
[0063] As an optional implementation, after migrating the file to be migrated after format conversion to the target table, the method further includes: determining the duration of format conversion of the file to be migrated as the current conversion duration; determining the duration of migrating the file to be migrated after format conversion to the target table as the current migration duration; determining the sum of the current conversion duration and the current migration duration as the current conversion migration duration; and updating the target mapping relationship according to the target conversion parameter value and the current conversion migration duration.
[0064] The current conversion duration may be the actual duration of the current format conversion operation on the to-be-migrated file. The current migration duration may be the actual duration of migrating the to-be-migrated file after format conversion in the source table to the target table. The current conversion and migration duration may be the sum of the current conversion duration and the current migration duration.
[0065] After migrating the format-converted file to be migrated to the target table, the actual duration of the format conversion operation on the file to be migrated and the actual duration of migrating the format-converted file to be migrated from the source table to the target table are respectively obtained to determine the actual total duration of the current file migration (i.e., the current conversion and migration duration). The target mapping relationship is then updated based on the current conversion and migration duration and the data volume of the file to be migrated processed this time (or currently), thereby achieving real-time dynamic updating of the target mapping relationship, reducing the risk of errors and data loss, and ensuring data accuracy and integrity.
[0066] Specifically, the target mapping relationship is updated according to the target conversion parameter value and the current conversion migration duration, including: when the current conversion migration duration is included in the multiple conversion migration durations corresponding to the target data volume interval, the conversion parameter value corresponding to the current conversion migration duration in the target mapping relationship is updated to the target conversion parameter value; when the current conversion migration duration is not included in the multiple conversion migration durations corresponding to the target data volume interval, an association relationship between the current conversion migration duration, the target conversion parameter value and the target data volume interval is established in the target mapping relationship.
[0067] For the case where the current conversion migration duration is included in the multiple conversion migration durations corresponding to the target data volume interval where the data volume of the file to be migrated is located, it is only necessary to update the conversion parameter value corresponding to the conversion migration duration, and update the original recorded conversion parameter value to the target conversion parameter value; for the case where the current conversion migration duration is included in the multiple conversion migration durations corresponding to the target data volume interval where the data volume of the file to be migrated is located, it is necessary to establish an association relationship between the target data volume interval where the data volume of the file to be migrated is located and the current conversion migration duration in the target mapping relationship, so as to realize the update of the target mapping relationship.
[0068] As an optional implementation, Figure 3 Flowchart of batch conversion of Hive tables into compressed formats according to an embodiment of the invention. Figure 3 As shown, taking the conversion of snappy compression format to zlib compression format as an example, the specific process is as follows:
[0069] S301: Revoke the write (insert) permission of the original table (i.e., source table). Specifically, the Hive admin user (i.e., a user with Hive database administrator privileges) can be used to revoke the write permission of the original table through the revoke method.
[0070] S302, obtain the original table information (i.e., basic information of the source table), which may include: Hive table permissions, HDFS file permissions; the HDFS data path and partition paths corresponding to the original table; the original table fields, attributes, and partitions; the number of files in each partition of the original table, the data size, etc.
[0071] S303: Based on the basic information of the original table (i.e., the HDFS data path and partition paths corresponding to the original table in step S302; the fields, attributes, and partitions of the original table), a new table (i.e., the target table) is created with the same fields and partition settings as the original table but with a different compression format. When extracting the fields and partition information of the original table, the DESCRIBE FORMATTED or SHOW CREATETABLE command can be used to view and obtain the complete definition of the table, including its fields, partitions, storage location, file format, and compression type. These field definitions and partition information can then be extracted or copied from the metadata of the original table. The new table name can be set to "original_table_ZLIB" by setting "orc.compress = ZLIB" in the TBLPROPERTIES property.
[0072] S304, obtain the Hive table permissions, HDFS file permissions, the number of files in each partition of the original table, the data size, etc. in step S302, where:
[0073] Obtain permission-related information. For Hive table permission information, use the set role admin and show granton table commands to obtain the current permissions held by the table creation user (that is, the Hive admin user). Query the permissions of the table to be converted using the Hive admin user. Use the revoke insert on library table from user command to revoke the write permission for the table user, ensuring that both the query permission and the write permission are revoked. For HDFS data path permission information, use the showcreate table command to obtain the location of the original table. Use the hdfs-dfs-ls / location command to obtain the write user information (that is, the HDFS file user permission information). Ensure that the write user has sufficient permissions and that the permissions after writing to the new table are consistent with the original table permissions.
[0074] Obtain information about the number of files and data size. Use the hdfs-dfs-count command to obtain the file number and the hdfs-dfs-du command to obtain the data size. Parameter settings (i.e., conversion parameter values) are obtained from the conversion information. The specific judgment rule prioritizes file data size. Specifically, the target data size range (i.e., the target data size range) is selected based on the current table's file data (i.e., the data size of the files to be migrated). Within this range (i.e., the target data size range), the parameters (i.e., the target conversion parameter values) with the shortest conversion time (i.e., the duration of the migration) are selected as the conversion parameters for the current table.
[0075] Finally, record the conversion information. Record the current table file number and data size, parameter settings, conversion time, and other information.
[0076] S305: Fill in the missing partitions. For the partition without data in the original table (i.e., the third target partition), recreate it in the new table (i.e., the fourth target partition).
[0077] S306, partition relocation. Relocate all partitions of the new table to the current data path of the original table. Specifically, after completing the creation of the new table and data conversion, modify the actual physical storage location of all partition data of the new table to the storage location of the original table data on HDFS.
[0078] S307, table relocation: relocate the new table location to the current original table data path.
[0079] S308: Move the original table. Move the original table's HDFS path files to the / tmp path. Specifically, move the original table's data files or directories stored on HDFS to the / tmp directory of HDFS. To move the data files, use the command hdfsdfs-mv / bigdata / ... / t_xxxxx / tmp / t_xxxxxx. To relocate the partitions, use the following command:
[0080] alter table t_xxx_zlib partition(xx___datemode='incr',xx___sdate='20180131',xx___edate='20180131')set location' / bigdata / ... / t_xxxxxx / xx___datemode=incr / xx___sdate=20180131 / xxx___edate=20180131'.
[0081] S309: Move the new table. Rename the new table path to the original table. Specifically, after completing the data conversion or format adjustment, use HDFS file system operations to rename the new table's data directory to the original table's data directory. This logically makes the new table "replace" the original table's location, but in reality, the new table's data and structure overwrite the original table's data and structure.
[0082] S310: Rename the original Hive table (i.e., source table) to a temporary table name. Use the ALTER TABLE RENAME TO method to rename the original Hive table name to a temporary table with the suffix _tmp.
[0083] S311: Rename the new Hive table (the target table) to the original table name. Use the ALTER TABLE RENAME TO method to rename the new Hive table to the original table name. Verify this using SHOW CREATE TABLE.
[0084] S312: Delete the original Hive table information. Use the DROP TABLE command to delete the temporary table.
[0085] S313: Delete the original table's HDFS file. Use the hdfs-dfs-rm-skipTrash method to skip the recycle bin and delete the original table's data path data.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0087] This embodiment also provides a device for migrating files, which is used to implement the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0088] Figure 4 is a structural block diagram of an apparatus for migrating files according to an embodiment of the present invention. Figure 4As shown, the device includes: a creation module 402, which is used to create a target table based on basic information of a source table, wherein the source table and the target table support different file formats for storage; a determination module 404, which is used to determine a target conversion parameter value in a target mapping relationship according to the data volume of the file to be migrated when migrating the file in the source table to the target table, wherein the target mapping relationship records the correspondence between the data volume of the file migrated within a historical time range and the conversion parameter value; a migration module 406, which is used to perform format conversion on the file to be migrated using the target conversion parameter value, and migrate the file to be migrated after format conversion to the target table.
[0089] In an exemplary embodiment, the device is also used to determine a target data volume interval corresponding to the data volume of the file to be migrated in the target mapping relationship, wherein the target mapping relationship includes at least one data volume interval, each of the data volume intervals corresponds to multiple conversion migration durations, and each conversion migration duration corresponds to one of the conversion parameter values; determine a target conversion migration duration from the multiple conversion migration durations corresponding to the target data volume interval; and determine the conversion parameter value corresponding to the target conversion migration duration as the target conversion parameter value.
[0090] In an exemplary embodiment, the device is also used to determine the minimum conversion migration duration among the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, determine the conversion migration duration ranked in the middle among the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or, in the case where the target data volume interval corresponds to N conversion migration durations, sort the N conversion migration durations to obtain the sorted N conversion migration durations; divide the target data volume interval into N sub-data volume intervals, and establish a correspondence between each sub-data volume interval and each sorted conversion migration duration in the order of the sorted N conversion migration durations; determine the sub-data volume interval in which the data volume of the file to be migrated is located as the target sub-data volume interval, and determine the conversion migration duration corresponding to the target sub-data volume interval as the target conversion migration duration.
[0091] In an exemplary embodiment, the device is further used to determine the data volume of files in each partition of the source table, wherein the source table includes multiple partitions; determine the partition in the source table with a non-zero data volume of files as the first target partition; and migrate the files to be migrated in the first target partition to the target table.
[0092] In an exemplary embodiment, the apparatus is further configured to create a second target partition in the target table; and migrate the files to be migrated in the first target partition to the second target partition, wherein the paths of the first target partition and the second target partition are the same.
[0093] In an exemplary embodiment, the device is further used to determine a partition in the source table whose file data volume is zero as a third target partition; and create a fourth target partition in the target table, wherein the fourth target partition is an empty partition and has the same path as the third target partition.
[0094] In an exemplary embodiment, the device is also used to determine the duration of the format conversion of the file to be migrated as the current conversion duration; determine the duration of migrating the file to be migrated after format conversion to the target table as the current migration duration; determine the sum of the current conversion duration and the current migration duration as the current conversion migration duration; and update the target mapping relationship based on the data volume of the file to be migrated, the target conversion parameter value, and the current conversion migration duration.
[0095] In an exemplary embodiment, the device is also used to update the conversion parameter value corresponding to the current conversion migration duration in the target mapping relationship to the target conversion parameter value when the current conversion migration duration is included in the multiple conversion migration durations corresponding to the target data volume interval; and to establish an association relationship between the current conversion migration duration, the target conversion parameter value and the target data volume interval in the target mapping relationship when the current conversion migration duration is not included in the multiple conversion migration durations corresponding to the target data volume interval.
[0096] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0097] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0098] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0099] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0100] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0101] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the steps of the method described in each embodiment of the present application when executed by a processor.
[0102] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0103] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0104] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for migrating files, characterized in that: include: Creating a target table based on basic information of a source table, wherein the source table and the target table support different file formats for storage; When migrating files from the source table to the target table, determining a target conversion parameter value in a target mapping relationship according to the data volume of the files to be migrated, wherein the target mapping relationship records a correspondence between the data volume of the files migrated within a historical time range and the conversion parameter value; The format of the file to be migrated is converted according to the target conversion parameter value, and the file to be migrated after the format conversion is migrated to the target table.
2. The method according to claim 1, characterized in that Determine the target conversion parameter values in the target mapping relationship based on the data volume of the files to be migrated, including: Determining a target data volume interval corresponding to the data volume of the to-be-migrated file in the target mapping relationship, wherein the target mapping relationship includes at least one data volume interval, each of the data volume intervals corresponds to a plurality of conversion migration durations, and each conversion migration duration corresponds to one of the conversion parameter values; Determining a target conversion migration duration from a plurality of conversion migration durations corresponding to the target data volume interval; The conversion parameter value corresponding to the target conversion migration duration is determined as the target conversion parameter value.
3. The method according to claim 2, characterized in that Determining a target conversion migration duration from a plurality of conversion migration durations corresponding to the target data volume interval includes: Determine the minimum conversion migration duration among the multiple conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or Determine the conversion migration duration that is in the middle of the plurality of conversion migration durations corresponding to the target data volume interval as the target conversion migration duration; or In the case where the target data volume interval corresponds to N of the conversion and migration durations, the N conversion and migration durations are sorted to obtain the sorted N conversion and migration durations; the target data volume interval is divided into N sub-data volume intervals, and in the order of the sorted N conversion and migration durations, a correspondence between each sub-data volume interval and each sorted conversion and migration duration is established; the sub-data volume interval in which the data volume of the file to be migrated is located is determined as the target sub-data volume interval, and the conversion and migration duration corresponding to the target sub-data volume interval is determined as the target conversion and migration duration.
4. The method according to claim 1, wherein Migrating the format-converted file to be migrated to the target table includes: Determining the data volume of files in each partition of the source table, wherein the source table includes a plurality of the partitions; Determine a partition in the source table where the data volume of the files is non-zero as a first target partition; Migrate the files to be migrated in the first target partition to the target table.
5. The method according to claim 4, characterized in that Migrating the to-be-migrated files in the first target partition to the target table includes: creating a second target partition in the target table; Migrate the files to be migrated in the first target partition to the second target partition, wherein the paths of the first target partition and the second target partition are the same.
6. The method according to claim 4 or 5, characterized in that The method further comprises: Determine a partition in the source table where the data volume of the files is zero as a third target partition; A fourth target partition is created in the target table, wherein the fourth target partition is an empty partition and has the same path as the third target partition.
7. The method according to claim 3, characterized in that After migrating the format-converted file to be migrated to the target table, the method further includes: Determining the duration of the format conversion of the file to be migrated as the current conversion duration; Determine the time it takes to migrate the format-converted file to the target table as the current migration time; Determine the sum of the current conversion duration and the current migration duration as the current conversion migration duration; The target mapping relationship is updated according to the target conversion parameter value and the current conversion migration duration.
8. The method according to claim 7, characterized in that Updating the target mapping relationship according to the target conversion parameter value and the current conversion migration duration includes: In a case where the current conversion migration duration is included in the plurality of conversion migration durations corresponding to the target data volume interval, updating the conversion parameter value corresponding to the current conversion migration duration in the target mapping relationship to the target conversion parameter value; When the current conversion migration duration is not included in the multiple conversion migration durations corresponding to the target data volume interval, an association relationship among the current conversion migration duration, the target conversion parameter value and the target data volume interval is established in the target mapping relationship.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 8.
Citation Information
Cited By
Data processing method, system and device
CN121280144A
Information retrieval method and system
CN121614604A