Big data offline processing method
The method simplifies Spark operations for big data offline processing by using a sh file to configure SQL statements, reducing complexity and costs while enhancing efficiency across various data processing tasks.
Patent Information
- Application Number
- CN202411507892.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-27
AI Technical Summary
Existing big data offline processing methods face challenges with high operational complexity, learning curves, and increased costs due to the need for users to learn Scala syntax and manage multiple database APIs, and they are limited in handling diverse data scenarios.
A simplified big data offline processing method using a sh file to specify a Spark framework, allowing users to configure and execute SQL statements without needing to learn Scala, thus reducing operational complexity and enhancing efficiency across various data processing needs.
Simplifies Spark operations, lowers learning and operational costs, and improves big data offline processing efficiency by allowing non-technical users to handle diverse data processing tasks effectively.
Smart Images

Figure CN120045542A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for offline big data processing. Background Art
[0002] Big data, or massive data, refers to data whose volume is so large that it cannot be captured, managed, processed, and organized into information that helps enterprise business decisions more actively through mainstream software tools within a reasonable time. Big data offline processing technology refers to the process of extracting data from the source data storage system, performing operations such as data cleaning, transformation, and aggregation during the big data processing, and then storing the processed data in the target data storage system. This process usually requires a large amount of computing resources and storage resources.
[0003] The existing general solutions for big data offline computing generally use the method of programming with Apache Spark. With the continuous expansion of the Spark ecosystem, users may face a relatively high learning curve and configuration complexity in actual applications. Since in the process of data integration and data migration, users need to not only understand how to program and use Spark, but also learn Scala syntax to write Spark programs. Moreover, in data migration, they need to learn and understand the APIs of different databases, which increases the difficulty of operating big data and also greatly increases the company's operating costs. On the other hand, since Apache Spark is a general big data processing component, it processes data too simply and cannot cope with various scenario requirements. It is impossible to customize and develop specifically for a certain type of big data production environment. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems of difficult operation and low processing efficiency in the existing big data offline processing method, and provide a big data offline processing method, which solves the requirements of data integration tasks, is simple to operate, improves the efficiency of big data offline processing, and can meet the needs of various data processing.
[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions: A big data offline processing method includes the following steps: S1: Start the main class, parse the initial parameters in the startup command, and obtain the template file to be processed and the values of the variable parameters in the SQL statement; S2: Parse the obtained template file, configure system parameters, obtain the SQL statement list and the file names of the SQL statements in the database file; S3: In the database file, search for xml tags with the same name to obtain the SQL statements to be searched; S4: Assemble the SQL statements into an SQL list, return the SQL list to the main class, and the main class replaces the variable parameters in the SQL statements with the values of the variable parameters in the SQL statements in the startup command, then execute the SQL statements to perform offline processing on big data.
[0006] Through the method of the present invention, the operation of Spark (a distributed data processing framework) can be simplified. Business personnel do not need to learn how to use Spark, nor do they need to learn Scala syntax. They can operate and process offline big data conveniently and simply, which improves the efficiency of offline big data processing, meets the needs of various data processing, and at the same time reduces the learning cost of data analysts using Spark and the operation cost of the company.
[0007] Preferably, specify the data ETL class of the distributed data processing framework inside the Hive database through a sh file, set the SQL control template xml file for starting the distributed data processing framework in the sh file, and at the same time set the SQL corpus file inside the SQL control template file, then start the distributed data processing framework and use the distributed data processing framework to perform offline processing on big data.
[0008] Preferably, the step S2 includes: parsing the obtained template file, configuring the application name of the distributed data processing framework, starting the distributed data processing framework, configuring the Hive database name, starting the Hive database, configuring steps, and obtaining the SQL statement list and the file name of the database file where the relevant SQL statements are located from the step in the steps.
[0009] Preferably, the step S4 includes: configuring the distributed streaming processing platform, executing the SQL statements, and sending all the execution information of the SQL statements to the distributed streaming processing platform during the execution process.
[0010] Preferably, the offline big data processing method is executed in the Hive database, and the data migration between the Hive database and other databases is also included during the offline big data processing: migrating the Hive database to the JDBC database, migrating the JDBC database to the Hive database, migrating the Hive database to the HBase database, and migrating the HBase database to the Hive database.
[0011] Preferably, the data migration process between the Hive databases includes setting the parameters of the JDBC database, the SQL statement for extracting data from the Hive database, and the target table name for storing in the JDBC database, and then migrating the Hive database to the JDBC database.
[0012] Preferably, the data migration process between the Hive databases includes setting the parameters of the JDBC database, the SQL statements for extracting data from the JDBC database, and the target table names for storing into the Hive database, and migrating the JDBC database to the Hive database.
[0013] Preferably, the data migration process between the Hive databases includes setting the source table names for extracting data from the Hive database and the target table names for storing into the HBASE database, and migrating the Hive database to the HBASE database.
[0014] Preferably, the data migration process between the Hive databases includes setting the source table names for extracting data from the HBASE database and the target table names for storing into the Jive database, and migrating the HBASE database to the Hive database.
[0015] Preferably, the data migration between the Hive database and the JDBC database includes meeting the prerequisite: a table with the same structure as MySQL has been established in the Hive database.
[0016] Therefore, the present invention has the following beneficial effects: By specifying the Spark framework through the sh startup file, business personnel do not need to rewrite Scala programs. They only need to simply configure the sh file to use Spark for big data offline processing, improving the efficiency of big data offline processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall step flow chart of the big data offline processing method in the Hive database of the present invention.
[0018] Figure 2 It is the schematic diagram of data flow of the big data offline processing method of the present invention.
[0019] Figure 3 It is the flow chart of migrating the Hive database to the JDBC database in the second embodiment.
[0020] Figure 4 It is the flow chart of migrating the JDBC database to the Hive database in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0021] The present invention will be further described in detail below in conjunction with the drawings and the specific embodiments: Embodiment 1: This embodiment provides a big data offline processing method, as Figure 1As shown in the figure, its operation process is as follows: Step 1, start the main class, parse the initial parameters in the startup command, and obtain the template file to be processed and the values of the variable parameters in the SQL statement; Step 2, parse the obtained template file, configure the system parameters, and obtain the SQL statement list and the file names of the SQL statements in the database file; Step 3, in the database file, search for xml tags with the same name to obtain the SQL statements to be searched; Step 4, assemble the SQL statements into an SQL list, return the SQL list to the main class, and the main class replaces the variable parameters in the SQL statement with the values of the variable parameters in the SQL statement in the startup command, and executes the SQL statement list to process big data.
[0022] When native spark solves problems of data processing, data integration, and data migration, it is necessary to use the scala language to write spark programs to implement the calculation operations of spark on hive data, the mutual migration between hive and jdbc databases, and the mutual migration between hive and hbase databases. Each big data offline processing task requires a scale transformation to operate spark, and the processing method is relatively complex and the processing efficiency is low.
[0023] The big data offline processing method provided in this embodiment forms a spark framework by specifying a series of pre-written scala classes through a sh startup file. It does not require business personnel to rewrite scala programs every time they process offline big data. They only need to simply configure the sh file to operate spark big data, improving the efficiency of big data offline processing. It is not necessary for business personnel to learn the usage of spark or the scala syntax to operate big data, which is convenient and concise.
[0024] Next, through specific examples and specific application scenarios, the technical solutions and technical effects of the present invention will be further described. The following examples are explanations of the present invention, and the present invention is not limited to the following examples.
[0025] As Figure 2 shown, a big data offline processing method provided in this embodiment includes data processing within the hive database and data migration between the hive database and other databases.
[0026] Specifically manifested as: I. Data operations within the hive database.
[0027] Specifically manifested as Figure 1 shown, including: Step (1): Start the main program, parse the initial parameters in the startup command, and obtain the template file to be processed and the values of the variable parameters in the SQL statement.
[0028] The initial parameters include specifying a template template file, variable parameter value 1 in SQL, variable parameter value 2 in SQL, ……, variable parameter value N in SQL.
[0029] The Template template file is a template file used to define how to process elements in an xml document.
[0030] The SQL statement is a Structured Query Language, which is a database query and programming language used to access data and query, update, and manage relational database systems.
[0031] Step (2): Parse the obtained template file, configure system parameters, obtain the SQL statement list, and the file name of the SQL statement in the database file; The parsed template template file can configure the spark application name, start spark, configure the hive database name, start the hive database name, configure steps. From the step in steps, the SQL statement list can be obtained, and the file name of the SQLLIST.xml type file where the relevant SQL statement is located can be known in the step.
[0032] In a database, steps usually refer to independent units that execute database tasks or operations. Each step may contain one or more SQL statements. These steps can help execute tasks more effectively and monitor and manage the execution of tasks.
[0033] The SQLLIST.xml type file is a database file mainly used to store and manage data, execute queries, insertions, updates, and deletions using SQL (Structured Query Language), and manage the database structure.
[0034] Step (3): In the database file, find xml tags with the same name to obtain the SQL statement to be searched.
[0035] According to the content of the step, we can find xml tags with the same name in the SQLLIST.xml type file. The content corresponding to this xml tag is the SQL statement that the step needs to search for.
[0036] Step (4): Assemble the SQL statements into an SQL list, return the SQL list to the main class, and the main class replaces the variable parameters in the SQL statements with the numerical values of the variable parameters in the startup command and executes the SQL statements.
[0037] Return the SQL statements to steps, and assemble them into a list of SQL statements according to steps. Return the list of SQL statements to the main class, and then the main class replaces all the variable parameters in the SQL statements with the values of the variable parameters in the SQL statements of the startup command.
[0038] Configure Kafka, and then execute the list of SQL statements. During the execution process, send all the execution information of the SQL statements to Kafka.
[0039] Kafka is a distributed streaming processing platform for high-throughput, persistent, and distributed data stream processing. The main functions of Kafka include message publishing and subscription, efficient data processing, and can be used as the basis for a data stream processing platform to process real-time data streams, such as event processing, real-time analysis, and training of machine learning models. It can also transfer data to a data warehouse for advanced analysis and reporting.
[0040] Second, the data migration process between the Hive database and other databases.
[0041] In this embodiment, a fixed sh file template and a Spark program written in Scala are used.
[0042] The data migration process between the Hive database and other databases includes: migrating the Hive database to the JDBC database, migrating the JDBC database to the Hive database, migrating Hive to the HBASE database, and migrating HBASE to the Hive database.
[0043] The Hive database is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading. This is a mechanism that can store, query, and analyze large-scale data stored in Hadoop. The Hive data warehouse tool can map structured data files into a database table and provide SQL query functions, and can transform SQL statements into MapReduce tasks for execution.
[0044] JDBC, that is, Java Database Connectivity, is an application programming interface that standardizes how client programs access databases, providing methods such as querying and updating data in databases.
[0045] The HBASE database is a highly reliable, high-performance, column-oriented, scalable distributed storage system, used for efficient random read and write operations on large-scale data sets.
[0046] Specifically: (1) Migrate the Hive database to the JDBC database.
[0047] Set the database address, username, and password of the JDBC database, the SQL statement for extracting data from the Hive database, and the name of the target table to be stored in the JDBC database. Then the task of migrating Hive data to the JDBC database can be completed.
[0048] (2) Migrate the JDBC to the Hive database.
[0049] Set the database address, username, and password of the JDBC database, the SQL statement for extracting data from the JDBC database, and the name of the target table to be stored in the Hive database. Then the task of migrating JDBC data to the Hive database can be completed.
[0050] (3) Migrate Hive to the HBASE database.
[0051] Set the name of the source table for extracting data from the Hive database and the name of the target table to be stored in the HBASE database. Then the task of migrating Hive data to the HBASE database can be completed.
[0052] (4) Migrate HBASE to the Hive database.
[0053] Set the name of the source table for extracting data from the HBASE database and the name of the target table to be stored in the Hive database. Then the task of migrating HBASE data to the Hive database can be completed.
[0054] A big data offline processing method provided in this embodiment can very well solve the needs of data analysis and data integration tasks of ordinary technical personnel, that is, data analysts and business personnel, according to local conditions. It greatly improves the efficiency of writing Spark programs from scratch to complete tasks. It realizes the advantages of one-time R & D of code and reuse everywhere. There is no need to rewrite the Spark program every time for a new task, improving the efficiency of big data offline processing.
[0055] Embodiment 2: This embodiment provides a big data offline processing method. On the basis of Embodiment 1, it is further optimized based on Apache Spark by bringing in specific application scenarios.
[0056] The solution of this embodiment is a Spark framework formed by specifying a series of pre-written Scala classes through a sh startup file. Business personnel do not need to rewrite Scala programs. They only need to simply configure the sh file to use and operate on Spark big data. Business personnel do not need to learn how to use Spark, nor do they need to learn Scala syntax, and they can operate and use big data, which is convenient and concise.
[0057] This embodiment simplifies Spark operations, reduces the learning cost for data analysts to use Spark, and reduces the company's operating costs. Two types are used to solve the problem of big data offline processing, specifically manifested as: I. Data operations within Hive.
[0058] Use the method of writing SQL within an xml configuration file, in conjunction with a fixed sh file template and Spark written in Scala.
[0059] Specifically: Specify the data ETL class of Spark inside the Hive database through the sh file. This data ETL class has been pre-written. And the SQL control template xml file required for this Spark startup can be set in the sh file. Set the required SQL corpus file (the SQL corpus file is a series of pre-written SQL statements nested in the xml file) inside the SQL control template file. After setting a series of SQL statements embedded in the xml file, Spark can be started for big data offline processing.
[0060] The sh file is a Shell script file, usually with the suffix.sh, which contains a series of commands executed by the operating system's command interpreter. ETL (Extract Transform Load) is a key component of data processing and integration. Its main function is to extract data from different data sources, perform necessary transformations, and then load the data into the target system or data warehouse.
[0061] Furthermore, setting the SQL statements includes: Step (1): Start Spark SQL, and start the script with parameters.
[0062] The parameter order is: specify the file name of the xml_template.xml type, variable parameter value 1 in SQL, variable parameter value 2 in SQL, variable parameter value in SQL,..., variable parameter value N in SQL.
[0063] The main class starts to run object Uniform Entry OOP, which is the entry point for the program to run.
[0064] Specifically: Use the spark startup script to start the main program object Uniform Entry OOP, and parse the parameters in the startup command. We can parse out the xml_template.xml type files that we need to process, as well as the values of the variable parameters in the SQL statements.
[0065] Object Uniform Entry OOP is the main class (now updated to Uniform Entry OOP_repartition), which is the entry point for the program to run. In the configured spark environment, hive environment, and kafka environment, execute the SQL statements configured in xml_template.xml.
[0066] At the same time, the main class directly parses the xml_template.xml file. Parsing the xml_template.xml file is mainly used to configure SQL statements, which involves the assignment of variable parameters in the SQL statements and the location of the specified corpus file where the specific SQL statements are located.
[0067] The xml_template.xml file is characterized by being able to extract the already written SQL statements from different corpus files. The advantage of doing this is that if the SQL statements we need already exist, then we don't need to repeat writing the same SQL statements again. We only need to know the path where the SQL statements exist, and we can take them out for use. This can improve the reuse rate and thus improve the efficiency of big data offline processing. By parsing the xml_template.xml file, we can find SQLLIST.xml files with different names at different locations.
[0068] That is, by parsing the xml_template.xml type file, we can configure the spark application name, start spark, configure the hive name, start hive, configure steps, obtain the SQL statement list from the step in steps, and know the file name of the SQLLIST.xml type file where the relevant SQL statements are located in the step.
[0069] Step (2): Parse KAFKAProperties.xml to obtain all kafka configurations.
[0070] The main class directly parses the KAFKAProperties.xml file, which configures the specific information of the kafka service. When spark SQL runs, all information needs to be configured and sent to the specified kafka.
[0071] Step (3): Through the parameters in the startup command, obtain the specified xml_template.xml file name, parse the xml_template.xml type file, and obtain the spark application name, hive database name, and step configuration to be run.
[0072] Step (4): By parsing the steps, obtain the address of the SQLLIST.xml type file to be parsed and the SQL tags to be extracted.
[0073] The SQLLIST.xml type file is a SQL statement corpus that stores SQL statements with variable parameters. This file can be newly written or previously written. SQLLIST.xml does not have to be called this file name, there can be many copies, and it can be in various locations. The main class locates various SQLLIST.xml type files by parsing the path in xml_template.xml.
[0074] According to the content of the step, we can find the xml tag with the same name in the SQLLIST.xml type file. The content corresponding to this xml tag is the SQL statement that the step needs to find. Return the SQL statement to the steps and assemble it into a SQL list according to the steps.
[0075] Step (5): Parse the SQLLIST.xml type file to obtain the SQL statement with variable parameters, and replace the variable parameters in the SQL statement with the values of the variable parameters in the startup command.
[0076] Return the SQL list to the main class object UniformEntryOOP, and then this main class replaces all the variable parameters in the SQL statement with the values of the variable parameters in the SQL statement in the startup command.
[0077] Step (6): Assemble the SQL statements included in the steps in a series of steps into a SQL statement list. objectUniformEntryOOP receives the SQL statement list, object UniformEntryOOP configures kafka, and objectUniformEntryOOP starts spark.
[0078] Step (7): object UniformEntryOOP executes the SQL statement list using Spark and sends the execution information to Kafka.
[0079] Configure Kafka and then execute the SQL statement list. During the execution process, all the execution information of the SQL statements is sent to Kafka.
[0080] Step (8): Perform big data offline processing.
[0081] The process of using the constructed Spark framework for big data offline processing includes: The offline data processing flow of Spark generally can be divided into the following steps: (8.1) Data reading: Read data from data sources (such as HDFS, Apache Hive, CSV, JSON, etc.).
[0082] (8.2) Data cleaning and preprocessing: Clean the raw data to remove invalid, duplicate, or incorrect data.
[0083] (8.3) Data transformation: Obtain the required results by transforming the data, such as aggregation, grouping, filtering, etc.
[0084] (8.4) Model training / analysis: Analyze the data according to business requirements or train machine learning models.
[0085] 5. Result output: Output the processed data to the specified data storage system.
[0086] II. The data migration process between the Hive database and other databases.
[0087] When native Spark solves problems of data processing, data integration, and data migration, it is necessary to write Spark programs in Scala language to implement the calculation operations of Spark on Hive data, the mutual migration between the Hive and JDBC databases, and the mutual migration between the Hive and HBase databases.
[0088] In this embodiment, the data migration process between the Hive database and other databases includes: migrating the Hive database to the JDBC database, migrating the JDBC database to the Hive database, migrating the Hive database to the HBASE database, and migrating the HBASE database to the Hive database.
[0089] The specific process of data migration between each database is as follows: (1), Migrate the Hive database to the JDBC database.
[0090] As Figure 3 shown, specifically represented as: Determine whether the preconditions are met, that is: a table with the same structure as MySQL has been created in Hive.
[0091] Set configuration items: (1.1) Open the hive2mysql.sh template file (export the statistical data in the Hive table to MySQL), and set the path of the main class, that is, the entry program class (class cn.creaway.datamigration.service.Hive2JDBC).
[0092] (1.2) Set the database address and set the database password.
[0093] (1.3) Determine whether the replace mode in MySQL is enabled, whether to update with the same primary key and insert with different primary keys.
[0094] (1.4) Set the Hive data source table, the Hive data source table query SQL statement, set the jdbc target table name, set the jdbc driver, and the jdbc concurrency can be increased when the data volume is relatively large. Due to a bug in Spark itself, only the incremental migration mode is provided in this embodiment. It is necessary to create the target table in Hive before performing data migration. When the number of fields in the source table and the target table is inconsistent, it needs to be adjusted through SQL statements. That is, Append cannot be modified, otherwise unknown errors will occur.
[0095] (2) Migrate from jdbc to the Hive database.
[0096] As Figure 4 shown, specifically represented as: Determine whether the preconditions are met, that is: a table with the same structure as MySQL has been created in Hive.
[0097] Set configuration items: (2.1) Open the hive2mysql.sh template file and set the path of the main class, that is, the entry program class (class cn.creaway.datamigration.service.Hive2JDBC).
[0098] (2.2) The jar package where the entry program class is located, set the database address, and set the database password.
[0099] (2.3) Determine whether the replace mode in MySQL is enabled, whether to update with the same primary key and insert with different primary keys.
[0100] (2.4) Set the MySQL source table (for the MySQL source table, SQL can be written), and set the name of the Hive target table (for the Hive target table, only the table name can be written).
[0101] (3) Migrate Hive to the HBASE database.
[0102] Specifically, it means: Open the HIVE2HBASE.sh template file, set the name of the Hive source table (hiveTable hiveTableName), and set the name of the HBase target table (HBaseTable HBaseTableName).
[0103] (4) Migrate HBASE to the Hive database.
[0104] Specifically, it means: (4.1) Open the HIVE2HBASE.sh template file, set the name of the HBase source table (HBaseTable HBaseTableName) and set the name of the Hive target table (HiveTable hiveTableName).
[0105] (4.2) Set the number of re-partitions for downloading data from HBase (in this embodiment, the number of re-partitions is 400), which can reduce the block size of each partition; (4.3) Set the specified time range for downloading the target table from HBase (in this embodiment, it can be expressed as: specified time range '2024-01-01 00:00:00,2024-12-30 00:00:00'), set the number of entries scanned from HBase each time (in this embodiment, it is set to: number of HBase entries 150), and when downloading, the number of entries scanned from HBase each time.
[0106] The big data offline processing method provided by this embodiment operates on the Spark file by configuring the sh file. There is no need to rewrite Spark from scratch for each new big data offline processing task, which greatly simplifies the process of business personnel using big data and improves the efficiency of big data offline processing.
[0107] The above-described embodiments are only a preferred solution of the present invention, and do not impose any form of limitation on the present invention. There are other variations and modifications without exceeding the technical solutions described in the claims.
Claims
1. A method for offline processing of big data, characterized in that: include: S1: Start the main class, parse the initial parameters in the startup command, obtain the template file to be processed and the values of the variable parameters in the SQL statement; S2: Parse the obtained template file, configure system parameters, obtain the SQL statement list and the file name of the SQL statement in the database file; S3: Search for the XML tag with the same name in the database file and obtain the SQL statement to be searched; S4: Assemble the SQL statements into an SQL list, return the SQL list to the main class, and the main class replaces the variable parameters in the SQL statement with the values of the variable parameters in the SQL statement in the startup command, executes the SQL statement list, and processes the offline big data.
2. A method for offline processing of big data according to claim 1, characterized in that: The data ETL class of the distributed data processing framework in the hive database is specified through the sh file. The SQL control template xml file for starting the distributed data processing framework is set in the sh file. At the same time, the SQL corpus file is set in the SQL control template file to start the distributed data processing framework, and the distributed data processing framework is used for offline processing of big data.
3. A method for offline processing of big data according to claim 1, characterized in that: The step S2 includes: parsing the obtained template file, configuring the distributed data processing framework application name, starting the distributed data processing framework, configuring the hive database name, starting the hive database, configuring steps, and obtaining the SQL statement list and the file name of the database file where the relevant SQL statement is located from the step in the steps.
4. A method for offline processing of big data according to claim 1, 2 or 3, characterized in that: The step S4 includes: configuring a distributed streaming processing platform, executing SQL statements, and sending all execution information of the SQL statements to the distributed streaming processing platform during the execution process.
5. The method for offline processing of big data according to claim 1, characterized in that: The big data offline processing method is executed in a hive database, and the big data offline processing process also includes data migration between the hive database and other databases: the hive database is migrated to the jdbc database, the jdbc database is migrated to the hive database, the hive database is migrated to the HBASE database, and the HBASE database is migrated to the hive database.
6. A method for offline processing of big data according to claim 5, characterized in that: The data migration process between the hive databases includes setting the parameters of the jdbc database, extracting the SQL statement of data from the hive database and the target table name stored in the jdbc database, and migrating the hive database to the jdbc database.
7. A method for offline processing of big data according to claim 5, characterized in that: The data migration process between the hive databases includes setting the parameters of the jdbc database, extracting the SQL statement of data from the jdbc database and the target table name stored in the hive database, and migrating the jdbc database to the hive database.
8. A method for offline processing of big data according to claim 5, 6 or 7, characterized in that: The data migration process between the hive databases includes setting the source table name for extracting data from the hive database and the target table name for storing data in the HBASE database, and migrating the hive database to the HBASE database.
9. A method for offline processing of big data according to claim 5, 6 or 7, characterized in that: The data migration process between the hive databases includes setting the source table name for extracting data from the HBASE database and the target table name for storing data in the hive database, and migrating the HBASE database to the hive database.
10. A method for offline processing of big data according to claim 5, 6 or 7, characterized in that: Data migration between the hive database and the jdbc database includes meeting the prerequisite: MySQL tables with the same structure have been created in the hive database.