Automated Loading Method and System for Big Data
Through automated loading methods and systems, using SparkSQL and MapReduce technologies, efficient and fault-tolerant big data migration is achieved, and multiple file formats are supported, solving the problem of inefficient data migration in the existing technology.
Patent Information
- Application Number
- CN202211556833.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-06
AI Technical Summary
In the existing technology, big data migration methods mainly rely on manual loading or own reading functions, and lack automated and efficient data migration methods, resulting in inefficient data mining.
The automated loading method is adopted, including data inspection and uploading, configuring data slicing and automatic reading steps, using SparkSQL and MapReduce technologies for data processing and slicing, and automatically reading data to the distributed file system HDFS through Hive, supporting multiple file formats.
It improves the speed and fault tolerance of big data loading, reduces data skew, supports multiple file formats, and enhances the universal adaptability of the system.
Smart Images

Figure CN115934189B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data migration and loading, and specifically, to a loading method and system for automatically loading big data. Background Art
[0002] With the continuous increase of the total national GDP, the banking industry has also witnessed rapid development, with the business volume continuously increasing, not only the improvement of offline business, but also the improvement of online business on both PC and mobile terminals; this has led to a significant increase in the data volume of different systems.
[0003] However, each business department is relatively independent. To deeply explore the value of different data, it is not limited to the data of a single department or a small number of departments. It is necessary to pool all the data of each business department in the whole company for in-depth exploration to improve the value of the data.
[0004] In order to securely store and efficiently explore and use data, migrating different data to a big data platform (Hadoop) has become a trend;
[0005] Patent document CN114296633A discloses a data migration method based on big data, including: the terminal simultaneously sends data acquisition requests to multiple metadata nodes; the first metadata node acquires the metadata corresponding to the file data in the first disk based on the data acquisition request, and sends the metadata corresponding to the file data to zookeeper; zookeeper acquires the first file data node corresponding to the file data and the corresponding second disk information based on the metadata corresponding to the file data; when the indication status of the second disk is saturated, zookeeper controls the second disk to migrate the data in the second disk to the third disk; zookeeper sends the first file data node and the mounted third disk information to the first metadata node, and forwards them to the terminal through the first metadata node.
[0006] However, the current data migration methods are usually to manually load data or use the built-in reading function of big data components to read data. Summary of the Invention
[0007] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a loading method and system for automatically loading big data.
[0008] According to a loading method for automatically loading big data provided by the present invention, it includes:
[0009] Data inspection and upload step: Inspect and upload the data file and the data flag file to the distributed file system HDFS;
[0010] Configuration data slicing step: Perform initialization configuration, process the data produced in the data check and upload step, slice it, and write it to the buffer areas on different machines in the cluster;
[0011] Automatic reading step: Write the data in the buffer area to the distributed file system HDFS to form a new file, and automatically read the new file written to the distributed file system HDFS directory through hive.
[0012] Preferably, the data check and upload step includes:
[0013] Step S1: Check the data file and the data flag file. If the number of data rows described in the data flag file is not 0 and the size of the data file described is the same as the actual size of the data file, then trigger Step S2; wherein, the data file is the storage document of the data to be loaded, and the data flag file is the basic description of the data file; in Step S1: If the number of data rows in the data flag file is 0, then do not load the data and the data loading is completed; if the size of the data file in the data flag file is different from the actual size of the data file, then it is considered abnormal and the user is prompted to check the data source;
[0014] Step S2: Upload the data file and the data flag file to the distributed file system HDFS.
[0015] Preferably, the configuration data slicing step includes:
[0016] Step S3: Register SparkSQL and enable HiveSupport, read the field information in the data flag file through SparkSQL, and transform the field information into a hive table creation statement, execute SparkSQL to create the corresponding hive table; at the same time, output the table structure statement as a json file to the HDFS temporary file;
[0017] Step S4: Initialize the basic information of the loading method;
[0018] Step S5: Establish a serialized object, and give different numbered codes for different field types; read the temporary file generated in Step S3, and put the generated table structure information into the object. At this time, the object obtains the starting position, ending position, data type, and maximum data occupancy length of each field in bytes;
[0019] Step S6: Configure the basic information of the MapReduce Job;
[0020] Step S7: Rewrite the MapReduce slicing mechanism, slice the data processed in Step S5 to obtain sliced data;
[0021] Step S8: Perform the step of the map() function on the sliced data and evenly write it to the buffers on different machines in the cluster;
[0022] Among them, the slicing methods include specifying a line break character, denoted as Method 1, and also include slicing according to the Byte length of the fields, denoted as Method 2; the primary key of the output data of Method 1 is of the IntWriteable type, and the value is of the Text type. The primary key of the output data of Method 2 is of the IntWriteable type, and the value is of the ByteWriteable type; among them, by default, data is loaded based on Method 1. When there is dirty data in the data, if the number of loaded rows is inconsistent with the data file, data is automatically loaded based on Method 2.
[0023] Preferably, the automatic reading step includes:
[0024] Step S9: When the data cache quantity exceeds 100M or after the map() function finishes processing all the data, the data undergoes a reduce() function operation, writes the data to HDFS, and assigns a unique serial number value to each data before writing. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function. Hive will automatically read the new files written to the HDFS directory;
[0025] Step 10: SparkSQL checks the amount of data imported into the hive table, compares the obtained data amount with the files in the flg file information. If the number of data rows is consistent, it is considered a successful load; otherwise, it is considered a failed load.
[0026] Preferably, data is automatically loaded in the way of MapReduce. Among them, when new fields are added after the original fields in the data source, the data loaded by reading the table creation statement still loads the data used in the original business.
[0027] According to an automated loading system for big data provided by the present invention, it includes:
[0028] Data check and upload module: Check and upload the data file and the data flag file to the distributed file system HDFS;
[0029] Configured data slicing module: Perform initialization configuration, slice the data processed by the data check and upload module, and write it to the buffers on different machines in the cluster;
[0030] Automatic reading module: Write the data in the buffer to the distributed file system HDFS to form new files, and hive automatically reads the new files written to the distributed file system HDFS directory.
[0031] Preferably, the data check and upload module includes:
[0032] Module M1: Check the data file and the data flag file, and if the number of data lines described in the data flag file is not 0 and the size of the data file described is the same as the actual size of the data file, trigger Module M2; wherein, the data file is a storage document of the data to be loaded, and the data flag file is a basic description of the data file; in Module M1: if the number of data lines in the data flag file is 0, then do not load the data and the data loading is completed; if the size of the data file in the data flag file is different from the actual size of the data file, it is considered abnormal and the user is prompted to check the data source;
[0033] Module M2: Upload the data file and the data flag file to the distributed file system HDFS.
[0034] Preferably, the configuration data slicing module includes:
[0035] Module M3: Register SparkSQL and enable HiveSupport, read the field information in the data flag file through SparkSQL, and convert the field information into a hive table creation statement, execute SparkSQL to create the corresponding hive table; at the same time, output the table structure statement as a json file to the HDFS temporary file;
[0036] Module M4: Initialize and load the basic information of the system;
[0037] Module M5: Establish a serialized object, and give different number codes for different field types; read the temporary file generated in Module M3, and put the generated table structure information into the object. At this time, the object obtains the starting position, ending position, data type, and maximum data occupation length of each field in bytes;
[0038] Module M6: Configure the basic information of the MapReduce Job;
[0039] Module M7: Rewrite the MapReduce slicing mechanism, slice the data processed by Module M5 to obtain sliced data;
[0040] Module M8: A module for the map() function on the sliced data, and evenly write it into the buffer areas on different machines in the cluster;
[0041] Among them, the slicing methods include specifying a line break character, denoted as Method 1, and also include slicing according to the Byte length of the field, denoted as Method 2; the primary key of the output data of Method 1 is of the IntWriteable type, and the value is of the Text type, and the primary key of the output data of Method 2 is of the IntWriteable type, and the value is of the ByteWriteable type.
[0042] Preferably, the automatic reading module includes:
[0043] Module M9: When the data cache quantity exceeds 100M or after the map() function finishes processing all the data, the data is operated by the reduce() function, written to HDFS, and a unique serial number value is added to each data before writing. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function, and hive will automatically read the new files written to the HDFS directory;
[0044] Module 10: SparkSQL checks the amount of data imported into the hive table, compares the obtained data amount with the files in the flg file information. If the number of data rows is the same, it is considered a successful load; otherwise, it is considered a failed load.
[0045] Preferably, the data is automatically loaded by means of MapReduce. Among them, when new fields are added after the original fields in the data source, the data loaded by reading the table creation statement still loads the data used in the original business.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. The present invention automatically loads data by means of MapReduce. Among them, when new fields are added after the original fields in the data source, the data loaded by reading the table creation statement still loads the data used in the original business, increasing the fault tolerance of data loading.
[0048] 2. The present invention writes by rewriting the slicing method, reducing data skew and thus improving the loading efficiency.
[0049] 3. The present invention loads data according to the default line break character, improving the efficiency of data loading. When there is dirty data in the data, if the number of loaded rows is inconsistent with the data file, it automatically executes loading according to the byte method, improving the loading accuracy. Different configurations of the output format can generate HDFS file formats such as TexFile, Parquet, and ORC, supporting the generation of mainstream HDFS document formats to meet different business requirements. Description of the Drawings
[0050] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non - limiting embodiments with reference to the accompanying drawings:
[0051] Figure 1 Schematic diagram of the process steps for providing a loading method for the invention.
[0052] Figure 2 Schematic diagram of the process steps for processing data files and data flag files in the present invention. Detailed implementation manners
[0053] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0054] The present invention improves the speed of big data automated loading, improves the fault tolerance of mapReduce, reduces the possibility of data skew, and basically covers the requirements of different formats such as textFile, Parquet, ORC, etc. after big data is read in, making the automated loading system more generally adaptable. It includes: obtaining, checking, and uploading data files and data flag files, reading file information to create hive table and json table structure files; initializing the file loading operation to create a file serialization storage object; rewriting steps such as slicing, map(), and Reduce() to implement the loading of files in MapReduce, and finally checking the data loaded into the table against the original data. Thus, a method for automatically loading big data files is realized, which improves the speed of big data automated loading, improves the fault tolerance of mapReduce, reduces the possibility of data skew, and basically covers the requirements of different formats such as textFile, Parquet, ORC, etc. after big data is read in, making the automated loading system more generally adaptable.
[0055] According to a loading method for automatically loading big data provided by the present invention, it includes the following steps:
[0056] Step 1: First, check the data file and the data flag file. The data flag file refers to the basic description of the data file, including the data file name, data file size, number of data rows in the data file, field information, etc. If the number of data rows in the data flag file is 0, do not load the data and the data loading is completed. If the data file size in the data flag file is different from the size of the data file, throw an exception to prompt to check the data source. If the data is not empty and the size of the data file is consistent with the data flag file information, then execute Step 2;
[0057] Step 2: Upload the obtained data file and data flag file to HDFS (Hadoop Distributed File System);
[0058] Step 3: Register SparkSQL and enable HiveSupport. Read the field information in the data flag file through SparkSQL, and transform the field information into a hive table creation statement, and execute SparkSQL to create the corresponding hive table. At the same time, output the table structure statement as a json file to the HDFS temporary file;
[0059] Step 4: Initialize the basic information of the loading method, including: input path, encoding format, basic configuration file of Hadoop, output file, output path;
[0060] Step 5: Establish a serialized object, and give different numbered codes for different field types. The codes are as follows:
[0061] 1: String, 2: decimal, 3: int, 4: bigint, 5: double, 6: float, 7: boolean, 8: date, etc., including common data types. Read the file generated in Step 3, and put the generated table structure information into the object. At this time, the object obtains the start position, end position, data type, and maximum data occupation length of each field in bytes;
[0062] Step 6: Configure the basic information of the MapReduce Job, including: configuration of the table structure information in Step 5, Job name, setting of the main execution class name;
[0063] Step 7: Rewrite the MapReduce slicing mechanism to slice the data. The slicing method can be specified as method 1 with the line break character, or sliced according to the Byte length of the field as method 2;
[0064] Step 8: Perform the step of the map() function on the data. Write the segmented data evenly into the buffers on different machines in the cluster through random numbers. The primary key of the output data in Method 1 is of the IntWriteable type, and the value is of the Text type. While the primary key of the output data in Method 2 is of the IntWriteable type, and the value is of the ByteWriteable type;
[0065] Step 9: When the data buffer quantity exceeds 100M or after the map() function finishes processing all the data, perform the reduce() function operation on the data, write the data to HDFS, and assign a unique serial number value to each data before writing. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function. Hive will automatically read the new files written to the HDFS directory;
[0066] Step 10: SparkSQL checks the amount of data imported into the hive table, compares the obtained data amount with the files in the flg file information. If the number of data rows is the same, the loading is successful; otherwise, the loading fails.
[0067] The present invention also provides a loading system for automatically loading big data. Those skilled in the art can implement the loading system for automatically loading big data by executing the process steps of the method for automatically loading big data. That is, the method for automatically loading big data can be understood as the preferred implementation manner of the loading system for automatically loading big data. Specifically, the loading system for automatically loading big data includes:
[0068] Data inspection and upload module: Inspect and upload the data file and data flag file to the distributed file system HDFS;
[0069] Configured data slicing module: Perform initialization configuration, slice the data processed by the data inspection and upload module, and write it into the buffers on different machines in the cluster;
[0070] Automatic reading module: Write the data in the buffer to the distributed file system HDFS to form new files, and let hive automatically read the new files written to the distributed file system HDFS directory.
[0071] The data inspection and upload module includes:
[0072] Module M1: Check the data file and the data flag file. If the number of data lines described in the data flag file is not 0 and the size of the data file described is the same as the actual size of the data file, trigger Module M2. Here, the data file is a storage document for the data to be loaded, and the data flag file is a basic description of the data file. In Module M1: If the number of data lines in the data flag file is 0, do not load the data and the data loading is completed. If the size of the data file in the data flag file is different from the actual size of the data file, it is considered an exception and the user is prompted to check the data source.
[0073] Module M2: Upload the data file and the data flag file to the distributed file system HDFS.
[0074] The configuration data slicing module includes:
[0075] Module M3: Register SparkSQL and enable HiveSupport. Read the field information in the data flag file through SparkSQL, and transform the field information into a hive table creation statement, and execute SparkSQL to create the corresponding hive table. At the same time, output the table structure statement as a json file to the HDFS temporary file.
[0076] Module M4: Initialize and load the basic information of the system.
[0077] Module M5: Establish a serialized object, and assign different numbered codes to different field types. Read the temporary file generated in Module M3, and put the generated table structure information into the object. At this time, the object obtains the starting position, ending position, data type, and maximum data occupancy length of each field in bytes.
[0078] Module M6: Configure the basic information of the MapReduce Job.
[0079] Module M7: Rewrite the MapReduce slicing mechanism, slice the data processed by Module M5 to obtain sliced data.
[0080] Module M8: A module for the map() function on the sliced data, and evenly write it into the buffer areas on different machines in the cluster.
[0081] Among them, the slicing methods include specifying a line break character, denoted as Method 1, and also include cutting according to the Byte length of the field, denoted as Method 2. The primary key of the output data of Method 1 is of the IntWriteable type, and the value is of the Text type. The primary key of the output data of Method 2 is of the IntWriteable type, and the value is of the ByteWriteable type.
[0082] The automatic reading module includes:
[0083] Module M9: When the data cache quantity exceeds 100M or after the map() function finishes processing all data, the data undergoes a reduce() function operation, and is written to HDFS. And before writing, a unique serial number value is added to each piece of data. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function. Hive will automatically read the new files written to the HDFS directory;
[0084] Module 10: SparkSQL checks the amount of data imported into the hive table, and compares the obtained data amount with the files in the flg file information. If the number of data rows is the same, it is considered a successful load; otherwise, it is considered a failed load.
[0085] Automatically load data in the way of MapReduce. Among them, when new fields are added after the original fields in the data source, the data is still loaded in the way of reading and loading through the table creation statement according to the data originally used in the business.
[0086] Those skilled in the art know that in addition to implementing the system, device and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system, device and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same program. Therefore, the system, device and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structures within the hardware component.
[0087] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. An automated method for loading big data, characterized in that, include: Data check and upload steps: Check and upload data files and data marker files to the distributed file system HDFS; Configure data slicing steps: Perform initial configuration, slice the data generated in the data inspection and upload steps, and write them to the cache areas on different machines in the cluster; Automatic reading step: Write the data in the cache area to the distributed file system HDFS to form a new file, and automatically read the new file written to the distributed file system HDFS directory through Hive; The step of configuring data slicing includes: Step S3: Register SparkSQL and enable HiveSupport. Use SparkSQL to read the field information in the data markup file, convert the field information into Hive table creation statements, execute SparkSQL to create the corresponding Hive table; and output the table structure statement as a JSON file to a temporary file on HDFS. Step S4: Initialize basic information of the loading method; Step S5: Create a serialized object and assign different number codes to different field types; read the temporary file generated in step S3, and put the generated table structure information into the object. At this time, the object obtains the starting position, ending position, data type, and maximum length of each field in Byte; Step S6: Configure the basic information of the MapReduce Job; Step S7: rewrite the MapReduce slicing mechanism to slice the data processed in step S5 to obtain sliced data; Step S8: Perform the map() function on the sliced data and write it evenly into the cache areas on different machines in the cluster; Among them, the slicing methods include specifying a line break, recorded as method 1, and also including cutting according to the Byte length of the field, recorded as method 2; the output data primary key of method 1 is of IntWriteable type, and the value is of Text type, and the output data primary key of method 2 is of IntWriteable type, and the value is of ByteWriteable type; among them, the data is loaded based on method 1 by default. When there is dirty data in the data, the number of loaded rows is inconsistent with the data file, and data loading based on method 2 is automatically executed.
2. The loading method for automatically loading big data according to claim 1, wherein, The data checking and uploading step includes: Step S1: Check the data file and the data flag file, and if the number of data rows described in the data flag file is not 0 and the size of the data file described is consistent with the actual size of the data file, then trigger step S2; wherein the data file is a storage document of the data to be loaded, and the data flag file is a basic description of the data file; in step S1: if the number of data rows in the data flag file is 0, then no data is loaded, and the data loading is completed; if the size of the data file in the data flag file is different from the actual size of the data file, it is considered an abnormality, and the user is prompted to check the data source; Step S2: Upload the data file and the data mark file to the distributed file system HDFS.
3. The loading method for automatically loading big data according to claim 1, characterized in that, The automatic reading step comprises: Step S9: When the data cache quantity exceeds 100M or after the map() function finishes processing all data, the data undergoes a reduce() function operation, is written to HDFS, and a unique serial number value is assigned to each piece of data before writing. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function, and Hive will automatically read the new files written to the HDFS directory; Step S10: SparkSQL checks the amount of data imported into the Hive table, compares the obtained data volume with the files in the flg file information. If the number of data rows is the same, it is considered a successful load; otherwise, it is considered a failed load.
4. The method for automatically loading big data according to claim 1, characterized in that: Automatically load the data through the MapReduce method. When new fields are added after the original fields in the data source, the data loaded by reading the table creation statement still loads the data used in the original business.
5. An automated big data loading system, characterized in that, Including: Data check and upload module: Check and upload the data file and data flag file to the distributed file system HDFS; Configured data slicing module: Perform initialization configuration, slice the data processed in the data check and upload module, and write it to the buffer areas on different machines in the cluster; Automatic reading module: Write the data in the buffer area to the distributed file system HDFS to form new files, and Hive automatically reads the new files written to the distributed file system HDFS directory; The configured data slicing module includes: Module M3: Register SparkSQL and enable HiveSupport, read the field information in the data flag file through SparkSQL, transform the field information into a Hive table creation statement, and execute SparkSQL to create the corresponding Hive table; at the same time, output the table structure statement as a json file to the HDFS temporary file; Module M4: Initialize the basic information of the loading system; Module M5: Create a serialized object and assign different number codes for different field types; read the temporary file generated in Module M3, put the generated table structure information into the object. At this time, the object obtains the starting position, ending position, data type, and maximum data occupation length of each field in bytes; Module M6: Configure the basic information of the MapReduce Job; Module M7: Rewrite the MapReduce slicing mechanism, slice the data processed in Module M5 to obtain sliced data; Module M8: A module that performs the map() function on the sliced data and evenly writes it to the buffer areas on different machines in the cluster; Among them, the slicing methods include specifying a line break character, denoted as Method 1, and also include cutting according to the Byte length of the field, denoted as Method 2; the output data primary key of Method 1 is of the IntWriteable type, and the value is of the Text type. The output data primary key of Method 2 is of the IntWriteable type, and the value is of the ByteWriteable type.
6. The loading system for automatically loading big data according to claim 5, wherein The data check and upload module includes: Module M1: Check the data file and the data flag file. If the number of data lines described in the data flag file is not zero and the size of the data file described is the same as the actual size of the data file, trigger Module M2. Herein, the data file is a storage document for the data to be loaded, and the data flag file is a basic description of the data file. In Module M1: If the number of data lines in the data flag file is zero, do not load the data and the data loading is completed. If the size of the data file in the data flag file is different from the actual size of the data file, it is considered an anomaly and the user is prompted to check the data source. Module M2: Upload the data file and the data flag file to the distributed file system HDFS.
7. The loading system for automatically loading big data according to claim 5, wherein The automatic reading module includes: Module M9: When the data cache quantity exceeds 100M or after the map() function finishes processing all the data, the data undergoes a reduce() function operation, and the data is written to HDFS. Before writing, a unique serial number value is assigned to each data. The data written to HDFS is output to the corresponding location according to the output path and data format configured by the reduce() function. Hive will automatically read the new files written to the HDFS directory. Module M10, SparkSQL checks the amount of data imported into the hive table, compares the obtained data volume with the files in the flg file information. If the number of data lines is consistent, it is considered a successful load; otherwise, it is considered a failed load.
8. The automatic big data loading system according to claim 5, characterized in that: Automatically load the data in the way of MapReduce. When new fields are added after the original fields in the data source, the data loaded by reading the table creation statement still loads the data used in the original business.
Citation Information
Patent Citations
Data migration method and system based on big data
CN114296633A
HBase loaded data importing method
CN103617211A
MapReduce overflow improvement method based on MPI-IO
CN114116293A