Big data migration method, big data migration system, chip and storage medium
By directly processing and storing data using data migration tools, the problem of long data migration time caused by traditional storage devices is solved, and efficient data migration is achieved.
Patent Information
- Application Number
- CN202510088518.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art requires the use of multiple traditional storage devices during data migration, resulting in long data migration time and low efficiency.
By reading the data in the preset first database, processing and storing it in the preset second database using the preset data migration tool, the processing and restoring of the data is directly realized, and the use of traditional storage devices is avoided.
Save a lot of time, improve data migration efficiency, and simplify the data migration process.
Smart Images

Figure CN120123315A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data migration. More specifically, it relates to a method for migrating big data, a system for migrating big data, a chip, and a computer-readable storage medium. Background Art
[0002] With the rapid development of the information age, the amount of data shows an explosive growth trend. When the current database cannot meet the needs of users, in order to expand the storage space, relevant personnel need to use multiple traditional storage devices, divide the data into multiple batches, download it into multiple traditional storage devices, and then download the data in each traditional storage device into the target database. This data migration method takes a lot of time. Summary of the Invention
[0003] Embodiments of this application provide a method for migrating big data, a system for migrating big data, a chip, and a computer-readable storage medium.
[0004] The method for migrating big data according to the embodiments of this application includes: reading data in a preset first database to obtain data to be processed; using a preset data migration tool to process the data to be processed to obtain data to be migrated; and using a preset data migration tool to store the data to be migrated into a preset second database.
[0005] The system for migrating big data according to the embodiments of this application includes a reading module and a processing module. The reading module is used to read data in a preset first database to obtain data to be processed; the processing module is used to use a preset data migration tool to process the data to be processed to obtain data to be migrated; and use a preset data migration tool to store the data to be migrated into a preset second database.
[0006] The chip according to the embodiments of this application includes a memory and a processor. The memory is configured to store a computer program. When the processor executes the computer program, the following method for migrating big data is implemented: reading data in a preset first database to obtain data to be processed; using a preset data migration tool to process the data to be processed to obtain data to be migrated; and using a preset data migration tool to store the data to be migrated into a preset second database.
[0007] The computer-readable storage medium according to the embodiments of this application stores a computer program thereon. When the program is executed by a processor, the following method for migrating big data is implemented: reading data in a preset first database to obtain data to be processed; using a preset data migration tool to process the data to be processed to obtain data to be migrated; and using a preset data migration tool to store the data to be migrated into a preset second database.
[0008] In the big data migration method, big data migration system, chip and computer-readable storage medium provided by this application, data to be processed is obtained by reading data in a preset first database, and then the data to be processed is processed by a preset data migration tool to obtain data to be migrated. Finally, the data to be migrated is stored in a preset second database by using the preset data migration tool. This application no longer needs to use traditional storage devices to store the data to be migrated, but directly realizes the processing and re-storage of the data to be migrated by using the preset data migration tool, saving a lot of time and improving the data migration efficiency.
[0009] Additional aspects and advantages of the embodiments of this application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of this application. Brief Description of the Drawings
[0010] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0011] Figure 1 is a flowchart of the big data migration method according to some embodiments of this application;
[0012] Figure 2 is a structural diagram of the big data migration system according to some embodiments of this application;
[0013] Figure 3 is a flowchart of using a preset data migration tool to process data to be processed to obtain data to be migrated in the big data migration method according to some embodiments of this application;
[0014] Figure 4 is a flowchart of using a preset data processing program in a preset data import tool to process the first type of data to be processed to obtain the first data to be migrated under a preset data scheduling platform in the big data migration method according to some embodiments of this application;
[0015] Figure 5 is a flowchart of using a preset data migration tool to store the data to be migrated in a preset second database in the big data migration method according to some embodiments of this application;
[0016] Figure 6 is a flowchart of using a preset data migration tool to process data to be processed to obtain data to be migrated in the big data migration method according to other embodiments of this application;
[0017] Figure 7It is a schematic flowchart of a process of using a preset data migration tool to process data to be processed and obtain data to be migrated in some other implementation manners of the present application;
[0018] Figure 8 It is a schematic flowchart of a process of using a preset data migration tool to store the data to be migrated into a preset second database in some other implementation manners of the present application;
[0019] Figure 9 It is a schematic flowchart of a process of using a preset data migration tool to store the data to be migrated into a preset second database in some other implementation manners of the present application;
[0020] Figure 10 It is a schematic structural diagram of a chip in some implementation manners of the present application;
[0021] Figure 11 It is a schematic diagram of the connection state between a computer-readable storage medium and a processor in some implementation manners of the present application.
[0022] Main element symbol description:
[0023] Chip 100;
[0024] Memory 101, processing unit 102;
[0025] Big data migration system 10;
[0026] Reading module 11; processing module 12; first database 13; second database 14
[0027] Processor 20;
[0028] Computer-readable storage medium 200; computer program 202. Specific implementation manners
[0029] The following details the implementation manners of the present application. The examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals always represent the same or similar elements or elements with the same or similar functions. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the implementation manners of the present application, and should not be construed as a limitation on the implementation manners of the present application.
[0030] With the rapid development of the information age, the amount of data shows an explosive growth trend. Taking the power industry as an example, in recent years, the power metering data of the power system has shown an explosive growth trend. These data include real-time information such as user power consumption, voltage, and current, as well as the operation status data of various power equipment. These data are of great significance to the operation, management, decision-making analysis, etc. of power enterprises. However, due to the extremely large amount of these data, the traditional data storage method, that is, the traditional database, has been difficult to meet the needs of modern enterprises. In the case where the current database cannot meet the needs of users, in order to expand the storage space, relevant personnel need to use multiple traditional storage devices, divide the data in the database into multiple batches, download it into multiple traditional storage devices, and then download the data in each traditional storage device into the target database. This data migration method takes a lot of time. How to solve the problem that the data migration process takes a lot of time has become a difficult problem that needs to be solved urgently by those skilled in the art. To solve this problem, this application provides a method for migrating big data (as Figure 1 shown), a big data migration system 10 (as Figure 2 shown), a chip (as Figure 10 shown), and a computer-readable storage medium 200 (as Figure 11 shown).
[0031] Please refer to Figure 1 and Figure 2 . The big data migration method according to the embodiments of this application includes:
[0032] 03: Read the data in the preset first database 13 to obtain the data to be processed;
[0033] 05: Use a preset data migration tool to process the data to be processed to obtain the data to be migrated; and
[0034] 07: Use a preset data migration tool to store the data to be migrated into the preset second database 14.
[0035] The above big data migration method can be applied to the big data migration system 10. The big data migration system 10 according to the embodiments of this application includes a reading module 11 and a processing module 12. The reading module 11 is used to read the data in the preset first database 13 to obtain the data to be processed; the processing module 12 is used to use a preset data migration tool to process the data to be processed to obtain the data to be migrated; and use a preset data migration tool to store the data to be migrated into the preset second database 14.
[0036] The method in step 03 belongs to the method of the reading module 11, and the methods in steps 05 and 07 belong to the method of the processing module 12. Specifically, the reading module 11 is used to execute step 03, and the processing module 12 is used to execute steps 05 and 07. From Figure 2 it can be seen that the big data migration system 10 may further include a preset first database 13 and a preset second database 14. Among them, the preset first database 13 refers to the database that was originally used to store data. During the big data migration process this time, the data in the first database 13 needs to be migrated to the preset second database 14. The preset second database 14 refers to a database that can store more data than the preset first database 13.
[0037] More specifically, the first database 13 can generally be a relational database (Relational Database Management System, RDBMS). A relational database refers to a database that uses the relational model to organize data. Relational databases store data in the form of rows and columns for easy viewing by users. This series of rows and columns in a relational database is called a table, and these tables make up the entire relational database. Users retrieve data from the database through queries, and a query is an execution code used to define certain areas in the database. The relational model can be simply understood as a two-dimensional table model, and a relational database is a data organization composed of two-dimensional tables and the relationships between them. Relational databases usually include various relational database management systems (such as MySQL, Oracle, MsSQL, PostgreSQL, and DB2, etc.). Among them, MySQL is a relational database management system. MySQL stores data in different tables instead of putting all data in a large warehouse, which increases speed and flexibility. At the same time, the SQL language used by MySQL is the most commonly used standardized language for accessing databases. Oracle refers to a relational database management system of Oracle Corporation, also known as Oracle Database. Oracle's system has good portability, is easy to use, and has strong functions. It is suitable for various large, medium, small, and microcomputer environments and is a high-efficiency, reliable, and high-throughput database solution. MsSQL (Microsoft SQL Server) is a relational database management system developed by Microsoft. MsSQL provides powerful management and processing capabilities for data. At the same time, MsSQL can support the Windows operating system, so it is widely used in various enterprise-level database environments. PostgreSQL is an object-relational database management system based on the POSTGRES version 4.2 developed by the Department of Computer Science at the University of California, Berkeley. PostgreSQL supports most SQL standards and provides many other modern features, such as complex queries, foreign keys, triggers, views, transaction integrity, multi-version concurrency control, etc. Similarly, PostgreSQL can also be extended in many ways, such as by adding new data types, functions, operators, aggregate functions, index methods, procedural languages, etc. In addition, because the PostgreSQL license is very flexible, any user can use, modify, and distribute PostgreSQL for any purpose for free. Therefore, PostgreSQL is also widely used in various industries.DB2 is a relational database management system developed by IBM in the United States. Its main operating environments include UNIX (including IBM's own AIX), Linux, IBM i (also known as OS / 400), z / OS, and Windows server versions. DB2 is mainly applied to various large-scale application systems. It has good scalability and can support environments from mainframes to single-user scenarios, being applicable under all common server operating system platforms. DB2 provides high-level data availability, integrity, security, recoverability, and the ability to execute applications from small-scale to large-scale. DB2 adopts data grading technology, which enables mainframe data to be easily downloaded to LAN database servers, allowing client / server users and LAN-based applications to access mainframe data and making database localization and remote connection transparent. DB2 is famous for having a very complete query optimizer. Its outer joins improve query performance and support multi-task parallel queries. DB2 has good network support capabilities. Each subsystem can connect hundreds of thousands of distributed users and can activate thousands of active threads simultaneously, making it particularly suitable for large-scale distributed application systems.
[0038] However, although these relational databases provide powerful functions in data management, they also have some obvious disadvantages. Firstly, when dealing with a large amount of data or complex queries, relational databases may face performance bottlenecks. Especially during the storage and reading of big data, especially when complex joins and aggregations are required, these traditional relational databases are inefficient in read and write operations. Secondly, relational databases are relatively weak in horizontal scalability. Although strategies such as replication, partitioning, and sharding can improve performance, this often increases management complexity and may cause other data management problems in the entire database. Thirdly, when relevant personnel need to migrate data from one relational database to another or integrate multiple data sources into one relational database, it will consume a large amount of time. Fourthly, relational databases are mainly designed to handle structured data. For unstructured or semi-structured data, such as data in XML or JSON formats, or images, videos, etc., additional processing means need to be added to relational databases.
[0039] To this end, the preset second database 14 used in this application is usually a distributed database. Compared with relational databases, distributed databases have more advantages. Specifically, first, distributed databases support data parallel processing, that is, the same operation can be executed on multiple nodes simultaneously, which greatly improves the processing speed and performance. This is particularly important for real-time applications that require quick responses. Second, distributed databases improve performance and capacity through horizontal scaling (i.e., adding more nodes) rather than vertical scaling (i.e., increasing the hardware resources of a single node). This means that when dealing with massive big data, distributed databases can easily handle the increase in data volume by adding more server nodes without the need to upgrade the hardware configuration of a single server, thus reducing costs and complexity. Third, due to the strong scalability of distributed databases, if distributed databases are used to store data, relevant personnel no longer need to perform frequent data migration tasks. Fourth, distributed databases can better support unstructured and semi-structured data, such as documents, pictures, videos, etc., which gives them greater flexibility when dealing with complex and diverse data sets. Therefore, through the comparison between relational databases and distributed databases, relational databases are more suitable for storing big data. The specific steps of this application and other hardware structures will be described in detail below.
[0040] Specifically, the big data migration system 10 refers to a system that moves data from one database to another. Among them, the big data migration process usually involves three key parts: data extraction, processing, and storage. In this application, the big data migration system 10 is mainly used to read the data in the preset first database 13, that is, extract the data to be processed, and process the data to be processed to obtain the data to be migrated that can be stored in the preset second database 14. The big data migration system 10 includes a reading module 11 and a processing module 12. The reading module 11 is a module in the big data migration system 10 that reads data from the source database, that is, the preset first database 13. The reading module 11 has the characteristics of high efficiency, reliability, and flexibility, and can adapt to different data sources and migration requirements. The processing module 12 is a module in the big data migration system 10 that is responsible for receiving the data to be processed obtained by the reading module 11 and performing necessary operations such as conversion, cleaning, and aggregation to ensure the correctness and availability of the data in the target database.
[0041] Specifically, in step 03, the reading module 11 reads the data stored in the preset first database 13 from the preset first database 13, so as to obtain the data to be processed. Among them, since the data formats supported by different databases are different, the data read by the reading module 11 still needs to be further processed. Therefore, these data are called "data to be processed". In step 05 and step 07, the processing module 12 uses the preset data migration tool to process the data to be processed, so as to obtain the data to be migrated that can be stored in the preset second database 14, and continues to use the preset data migration tool to store the data to be migrated in the preset second database 14, so as to complete the data migration. Among them, the preset data migration tool includes a variety of data processing tools, which can be used to clean, transform, and filter the data to be processed, so as to obtain the data to be migrated.
[0042] Since a distributed database usually includes a distributed file system (HDFS), a distributed data warehouse (Hive), and a columnar repository (HBase), the data migration process of big data will be described in more detail in three cases in turn, that is, first migrating the data to the distributed file system, first migrating the data to the distributed data warehouse, and first migrating the data to the columnar repository.
[0043] In some embodiments, please refer to Figure 2 and Figure 3 , the data to be migrated includes the first data to be migrated; the preset data migration tool includes a preset data import tool, and step 05 includes:
[0044] 051: Classify the data to be processed according to the preset data classification program to obtain the first type of data to be processed; and
[0045] 052: Under the preset data scheduling platform, use the preset data processing program in the preset data import tool to process the first type of data to be processed to obtain the first data to be migrated.
[0046] The above big data migration method can be applied to the big data migration system 10. The processing module 12 is also used to: classify the data to be processed according to the preset data classification program to obtain the first type of data to be processed; and under the preset data scheduling platform, use the preset data processing program in the preset data import tool to process the first type of data to be processed to obtain the first data to be migrated.
[0047] It can be understood that different types of data stored in the first database 13 need to be processed using different data migration tools. Specifically, taking the relevant data of the power system as an example, the preset data import tool is usually Sqoop. Sqoop is an open-source tool mainly used for data import and export between a Hadoop cluster (i.e., the computing framework of a distributed database) and a traditional relational database. Sqoop can solve the problem of large-scale data migration, enabling data engineers to easily extract data from a relational database and load it into the distributed file system corresponding to the Hadoop cluster, or export data from the distributed file system corresponding to the Hadoop cluster to a relational database. The processing module 12 classifies the data to be processed through a preset data classification program to obtain the first type of data to be processed. The first type of data to be processed is usually historical current-voltage curve data and offline data in an Oracle database. Then, the processing module 12 processes the first type of data to be processed through a preset data processing program (usually a MapReduce program) in Sqoop to obtain the first data to be migrated. More specifically, Sqoop executes the MapReduce program in the Hadoop cluster, and through the MapReduce program, operations such as splitting, partitioning, reading, and writing of the first type of data to be processed are performed, that is, operations such as cleaning, transforming, and filtering the first type of data to be processed are performed to obtain the first data to be migrated. The core principle of Sqoop is to connect to the first type of data to be processed through JDBC (Java language connecting to the database), generate a Java class based on the first type of data to be processed, that is, generate a new JAVA file, and parallelly process the first type of data to be processed through the MapReduce program, thereby achieving efficient data transmission. Sqoop can also receive shell commands or Java API commands input by relevant personnel through the client and convert these commands into corresponding MapReduce programs through a task translator.
[0048] Please refer to Figure 2 and Figure 4 , in some embodiments, the preset data processing program includes a preset data cleaning program, a preset data transformation program, and a preset data filtering program. Step 052 includes:
[0049] 0521: Under the preset data scheduling platform, use the preset data cleaning program to clean the first data to be processed to obtain the data to be transformed;
[0050] 0523: Under the preset data scheduling platform, use the preset data transformation program to transform the data to be transformed to obtain the data to be filtered; and
[0051] 0525: Under a preset data scheduling platform, use a preset data filtering program to filter the first data to be filtered to obtain the first data to be migrated.
[0052] The above-mentioned big data migration method can be applied to the big data migration system 10. The processing module 12 is further configured to: under a preset data scheduling platform, use a preset data cleaning program to clean the first data to be processed to obtain the data to be converted; under a preset data scheduling platform, use a preset data conversion program to convert the data to be converted to obtain the data to be filtered; and under a preset data scheduling platform, use a preset data filtering program to filter the first data to be filtered to obtain the first data to be migrated.
[0053] Specifically, the preset data scheduling platform provides an operating environment for the preset data import tool. In this application, DolphinScheduler is usually used as the core scheduling platform for collecting system big data, that is, the preset data scheduling platform. Still taking the relevant data of the power system as an example, the preset data scheduling platform is used to deploy a preset data processing program, that is, the MapReduce program, and configure the functions of the preset data processing program. If divided from the functions of the program, the preset data processing program includes a data cleaning program, a data conversion program, and a data filtering program. Among them, data cleaning refers to identifying and processing missing values, detecting and correcting error data, and removing duplicate data, etc. Data conversion refers to performing format transformation on the data to be processed and converting the data format of the data to be processed into other data formats. Data filtering refers to the process of screening out data items or data sets that meet the requirements according to certain conditions during the data processing process. The processing module 12 first uses a preset data cleaning program to clean the first data to be processed under the preset data scheduling platform to obtain the data to be converted, then uses a preset data conversion program to convert the data to be converted to obtain the data to be filtered, and finally uses a preset data filtering program to filter the first data to be filtered to obtain the first data to be migrated.
[0054] Specifically, in step 0521, the cleaning of the first data to be processed can also be performed by Spark. Spark is a general big data computing framework that can complete the cleaning of some complex data. At the same time, Spark also has the function of transforming and migrating data. In step 0523 and step 0525, the transformation and filtering of the first data to be processed can also be performed by Kettle. Kettle is an open-source ETL tool mainly used for the process of data extraction, transformation, and loading. Relevant personnel can build a data pipeline by dragging components, connecting lines, and configuring, and can complete operations such as data reading, data association, data filtering, data format conversion, calculation, statistics, modeling, mining, and output to different target databases without writing code. Kettle can connect to multiple data sources and is applicable to various application scenarios of data transformation and data filtering.
[0055] Please refer to Figure 2 and Figure 5 , in some embodiments, the preset data migration tool further includes a preset distributed message queue. The preset second database 14 includes a distributed file system (HDFS), a distributed data warehouse (Hive), and a columnar repository (HBase). Step 07 includes:
[0056] 071: Under the preset data scheduling platform, use the preset data import tool and the preset distributed message queue to store the first data to be migrated into the distributed file system; and
[0057] 072: According to the preset data distribution program, use the preset data import tool and the preset distributed message queue to store the first data to be migrated stored in the distributed file system into the distributed data warehouse and the columnar repository.
[0058] The above big data migration method can be applied to the big data migration system 10. The processing module 12 is further configured to: under the preset data scheduling platform, use the preset data import tool and the preset distributed message queue to store the first data to be migrated into the distributed file system; and according to the preset data distribution program, use the preset data import tool and the preset distributed message queue to store the first data to be migrated stored in the distributed file system into the distributed data warehouse and the columnar repository.
[0059] Specifically, the preset distributed message queue usually includes Kafka, which is a distributed publish-subscribe message system and is usually used to assist the preset data import tool, namely Sqoop, to achieve real-time transmission of the first data to be migrated. Since both the distributed data warehouse and the columnar repository are databases based on the distributed file system at the bottom layer, when the processing module 12 stores the first data to be migrated, it needs to use the preset data import tool and the preset distributed message queue under the preset data scheduling platform to store the first data to be migrated into the distributed file system, and then store the first data to be migrated stored in the distributed file system into the distributed data warehouse and the columnar repository according to the preset data import tool and the preset distributed message queue.
[0060] Please refer to Figure 2 and Figure 6 In some embodiments, the preset data migration tool includes a preset data synchronization tool, the data to be migrated includes the second data to be migrated, and step 05 further includes:
[0061] 053: Classify the data to be processed according to the preset data classification program to obtain the second type of data to be processed; and
[0062] 054: Under the preset data scheduling platform, use the preset data synchronization tool to process the second type of data to be processed to obtain the second data to be migrated.
[0063] The above data migration method for big data can be applied to the big data migration system 10. The processing module 12 is further configured to: classify the data to be processed according to the preset data classification program to obtain the second type of data to be processed; and under the preset data scheduling platform, use the preset data synchronization tool to process the second type of data to be processed to obtain the second data to be migrated.
[0064] It can be understood that different types of data stored in the first database 13 need to be processed using different data migration tools. Specifically, taking the relevant data of the power system as an example, the preset data synchronization tool is usually DataX, which is a tool for efficiently migrating data to different data storage systems, such as migrating from a relational database to a distributed database. DataX has strong data migration capabilities, stability, and ease of use, and can support data migration between multiple data sources and target data storage systems. In this application, the processing module 12 classifies the data to be processed according to the preset data classification program to obtain the second type of data to be processed, which is usually power data. Then, the processing module 12 processes the second type of data to be processed through DataX under the preset data scheduling platform to obtain the second data to be migrated.
[0065] Please combineFigure 2 , in some embodiments, the preset data synchronization tool includes a preset data splitting program, and the data splitting program runs based on a preset data synchronization algorithm framework. Step 054 includes:
[0066] 0541: Based on a preset data scheduling platform and a preset data synchronization algorithm framework, use the preset data splitting program to split the second type of data to be processed, and obtain second data to be migrated composed of multiple groups of second sub-data to be migrated.
[0067] The above-mentioned big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: based on a preset data scheduling platform and a preset data synchronization algorithm framework, use the preset data splitting program to split the second type of data to be processed, and obtain second data to be migrated composed of multiple groups of second sub-data to be migrated.
[0068] More specifically, taking the relevant data of the power system as an example, the preset data splitting program includes Reader and Writer plugins, and the preset data synchronization algorithm framework includes the Framework+Plugin design pattern algorithm framework. The processing module 12 uses the preset data splitting program in DataX and adopts the task sharding method to split the second type of data to be processed into multiple small pieces of data, that is, multiple groups of second sub-data to be migrated, so as to improve the data migration speed in the subsequent data migration process.
[0069] Please combine Figure 2 , in some embodiments, the preset data migration tool further includes a preset distributed message queue, and the preset second database 14 includes a distributed data warehouse (Hive). Step 07 further includes:
[0070] 073: Under the preset data scheduling platform, use the preset data synchronization tool and the preset distributed message queue to store the second data to be migrated into the distributed data warehouse.
[0071] The above-mentioned big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: under the preset data scheduling platform, use the preset data synchronization tool and the preset distributed message queue to store the second data to be migrated into the distributed data warehouse.
[0072] Specifically, since in step 0541, the second type of data to be processed has been split into multiple groups of second sub-data to be migrated, that is, the second data to be migrated is composed of multiple groups of second sub-data to be migrated. Therefore, in step 073, the processing module 12 uses the preset data synchronization tool and the preset distributed message queue under the preset data scheduling platform to store multiple groups of second sub-data to be migrated into the distributed data warehouse in parallel, thereby improving the data migration speed.
[0073] Please combine with Figure 2 , in some embodiments, the second database 14 includes a distributed data warehouse (Hive) and a distributed file system (HDFS), and the migration method of the present application further includes:
[0074] 074: Use a preset data storage program to read the second data to be migrated in the distributed data warehouse and store the second data to be migrated in the distributed file system.
[0075] The above-mentioned big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: use a preset data storage program to read the second data to be migrated in the distributed data warehouse and store the second data to be migrated in the distributed file system.
[0076] It can be understood that since the distributed data warehouse is a database based on a distributed file system at the bottom layer, the processing module 12 also needs to use a preset data storage program to store the second data to be migrated stored in the distributed data warehouse in the distributed file system.
[0077] Please refer to Figure 2 and Figure 7 , in some embodiments, the data to be migrated includes the third data to be migrated, and step 05 further includes:
[0078] 055: Classify the data to be processed according to a preset data classification program to obtain the third type of data to be processed;
[0079] 056: Receive a data processing instruction under a preset data scheduling platform; and
[0080] 057: Process the third type of data to be processed according to the data processing instruction under a preset data scheduling platform to obtain the third data to be migrated.
[0081] The above-mentioned big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: classify the data to be processed according to a preset data classification program to obtain the third type of data to be processed; receive a data processing instruction under a preset data scheduling platform; and process the third type of data to be processed according to the data processing instruction under a preset data scheduling platform to obtain the third data to be migrated.
[0082] It can be understood that in addition to processing the data to be processed according to various tools, the processing module 12 can also process the data to be processed according to the instructions input by relevant personnel. For example, the processing module 12 can receive the Shell commands input by relevant personnel under a preset data scheduling platform. The Shell command is a way of interaction between the operating system and the user, which allows the user to input instructions to execute system operations, manage files, run programs, view system information, control processes, etc. The processing module 12 processes the third type of data to be processed according to these Shell commands, so as to obtain the third data to be migrated.
[0083] Please refer to Figure 2 and Figure 8 , in some embodiments, the preset data migration tool includes a preset distributed message queue, the preset second database 14 includes a columnar storage repository (HBase), and step 07 further includes:
[0084] 075: Receive a first data migration instruction under a preset data scheduling platform; and
[0085] 076: In response to the first data migration instruction, use the preset distributed message queue to store the third data to be migrated into the columnar storage repository.
[0086] The above-mentioned big data migration method can be applied to the big data migration system 10. The processing module 12 is further configured to: receive a first data migration instruction under a preset data scheduling platform; and in response to the first data migration instruction, use the preset distributed message queue to store the third data to be migrated into the columnar storage repository.
[0087] For example, the first data migration instruction can be an import command. The processing module 12 can store the third data to be migrated into the columnar storage repository according to the import command and the preset distributed message queue under a preset data scheduling platform.
[0088] Please refer to Figure 2 and Figure 9 , in some embodiments, the preset second database 14 includes a target disk, and step 07 further includes:
[0089] 077: Receive a second data migration instruction under a preset data scheduling platform; and
[0090] 078: In response to the second data migration instruction, store the third data to be migrated into the target disk.
[0091] The above big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: receive a second data migration instruction under a preset data scheduling platform; and in response to the second data migration instruction, store the third data to be migrated into the target disk.
[0092] For example, the second data migration instruction may be an export command. The processing module 12 may directly store the data stored in the third data to be migrated into the target disk in the form of a data file according to the export command under the preset data scheduling platform.
[0093] Please refer to Figure 2 and Figure 9 , in some embodiments, the preset second database 14 includes a distributed file system (HDFS), and step 07 further includes:
[0094] 077: Receive a second data migration instruction under a preset data scheduling platform; and
[0095] 079: Store the third data to be migrated into the distributed file system according to the second data migration instruction.
[0096] The above big data migration method can be applied to the big data migration system 10, and the processing module 12 is further configured to: receive a second data migration instruction under a preset data scheduling platform; and store the third data to be migrated into the distributed file system according to the second data migration instruction.
[0097] For example, the second data migration instruction can be an export command. The processing module 12 can, under a preset data scheduling platform, according to the export command, directly store the data stored in the third data to be migrated into the distributed file system in the form of a data file. From step 078 and step 079, it can be seen that the processing module 12 can also directly store the data to be migrated according to the instructions input by relevant personnel. In addition, in addition to the instructions input by relevant personnel, the distributed database also stores programs, tools, or instructions for data import and export. For example, the columnar repository contains an Application Programming Interface (API). An application programming interface is a standard interface that defines how software components interact with each other. The application programming interface provides a series of predefined methods, functions, protocols, and data structures, enabling different software components to communicate and cooperate with each other without having to understand the internal implementation details of each other. Therefore, the application programming interface can be used for data import and export. For example, the put method of the HTable class can be used to import data into the HBase table, and the getScanner method of the Scan class can be used to export data to a file. And the distributed data warehouse can usually also migrate data by directly reading or copying data files. The relational database also has its own data export command, which can export the data stored in itself into data files such as csv, xml, and txt.
[0098] In summary, in the big data migration method provided in this application, by reading the data in the preset first database 13, the data to be processed is obtained, and then the preset data migration tool is used to process the data to be processed to obtain the data to be migrated. Finally, the preset data migration tool is used to store the data to be migrated into the preset second database 14. This application no longer needs to use traditional storage devices to store the data to be migrated, but uses a preset data migration tool to directly implement the processing and re-storage of the data to be migrated, saving a lot of time and improving the data migration efficiency.
[0099] Please refer to Figure 10 , in some embodiments, this application also provides a chip 100. The chip 100 includes a memory 101 and a processing unit 102. The memory 101 is configured to store a computer program. When the processing unit 102 executes the computer program, the method in any of the above embodiments is implemented.
[0100] For example, please combine Figure 2 , when the processing unit 102 executes the computer program stored in the memory 101, the following method is implemented:
[0101] 03: Read the data in the preset first database 13 to obtain the data to be processed;
[0102] 05: Use a preset data migration tool to process the data to be processed and obtain the data to be migrated; and
[0103] 07: Use a preset data migration tool to store the data to be migrated into a preset second database 14.
[0104] For another example, when the processing unit 102 executes the computer program stored in the memory 101, the following method is implemented:
[0105] 051: Classify the data to be processed according to a preset data classification program to obtain the first type of data to be processed; and
[0106] 052: Under a preset data scheduling platform, use a preset data processing program in a preset data import tool to process the first type of data to be processed to obtain the first data to be migrated.
[0107] For another example, when the processing unit 102 executes the computer program stored in the memory 101, the methods in 0521, 0523, 0525, 053, 054, 0541, 055, 056, 057, 071, 72, 073, 074, 075, 076, 077, 078, and 079 can also be implemented.
[0108] The chip 100 provided by this application reads the data in the preset first database 13 to obtain the data to be processed, then uses a preset data migration tool to process the data to be processed to obtain the data to be migrated, and finally uses a preset data migration tool to store the data to be migrated into the preset second database 14. This application no longer needs to use traditional storage devices to store the data to be migrated, but uses a preset data migration tool to directly implement the processing and re - storage of the data to be migrated, saving a large amount of time and improving the data migration efficiency.
[0109] Please refer to Figure 2 and Figure 11 , in some embodiments, this application also provides a computer - readable storage medium 200, on which a computer program 202 is stored, and when the program is executed by a processor, the method in any one of the above - mentioned embodiments is implemented.
[0110] For example, when the computer program 202 is executed by the processor 20, the following method is implemented:
[0111] 03: Read the data in the preset first database 13 to obtain the data to be processed;
[0112] 05: Use a preset data migration tool to process the data to be processed and obtain the data to be migrated; and
[0113] 07: Use a preset data migration tool to store the data to be migrated into a preset second database 14.
[0114] For another example, when the computer program 202 is executed by the processor 20, the following method is implemented:
[0115] 051: Classify the data to be processed according to a preset data classification program to obtain the first type of data to be processed; and
[0116] 052: Under a preset data scheduling platform, use a preset data processing program in a preset data import tool to process the first type of data to be processed to obtain the first data to be migrated.
[0117] For another example, when the computer program 202 is executed by the processor 20, the methods in 0521, 0523, 0525, 053, 054, 0541, 055, 056, 057, 071, 72, 073, 074, 075, 076, 077, 078, and 079 can also be implemented.
[0118] In the computer-readable storage medium 200 of the present application, the data to be processed is obtained by reading the data in the preset first database 13, and then a preset data migration tool is used to process the data to be processed to obtain the data to be migrated. Finally, a preset data migration tool is used to store the data to be migrated into the preset second database 14. The present application no longer needs to use a traditional storage device to store the data to be migrated, but uses a preset data migration tool to directly implement the processing and re-storage of the data to be migrated, saving a lot of time and improving the data migration efficiency.
[0119] In the description of this specification, the descriptions referring to terms such as "certain embodiments", "in an example", "exemplarily", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0120] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0121] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A big data migration method, characterized in that: The migration method comprises: Read the data in the preset first database to obtain the data to be processed; Using a preset data migration tool to process the data to be processed to obtain data to be migrated; and The data to be migrated is stored in a preset second database using a preset data migration tool.
2. The migration method according to claim 1, characterized in that: The data to be migrated includes first data to be migrated; the preset data migration tool includes a preset data import tool, and the using of the preset data migration tool to process the data to be processed to obtain the data to be migrated includes: Classifying the data to be processed according to a preset data classification program to obtain a first category of data to be processed; and Under the preset data scheduling platform, the first category of data to be processed is processed using a preset data processing program in the preset data import tool to obtain the first data to be migrated.
3. The migration method according to claim 2, characterized in that: The preset data processing program includes a preset data cleaning program, a preset data conversion program, and a preset data filtering program. Under the preset data scheduling platform, the preset data processing program in the preset data import tool is used to process the first type of data to be processed to obtain the first data to be migrated, including: Under the preset data scheduling platform, using the preset data cleaning program, cleaning the first data to be processed to obtain data to be converted; Under the preset data scheduling platform, using the preset data conversion program, converting the data to be converted to obtain data to be filtered; and Under the preset data scheduling platform, the preset data filtering program is used to filter the first data to be filtered to obtain the first data to be migrated.
4. The migration method according to claim 2, characterized in that: The preset data migration tool further includes a preset distributed message queue, the preset second database includes a distributed file system, a distributed data warehouse, and a column storage library, and the use of the preset data migration tool to store the data to be migrated into the preset second database includes: Under the preset data scheduling platform, using the preset data import tool and the preset distributed message queue, storing the first data to be migrated in the distributed file system; and According to a preset data distribution program, the preset data import tool and the preset distributed message queue are used to store the first data to be migrated stored in the distributed file system into the distributed data warehouse and the column storage library.
5. The migration method according to claim 1, characterized in that: The preset data migration tool includes a preset data synchronization tool, the data to be migrated includes second data to be migrated; and the using the preset data migration tool to process the data to be processed to obtain the data to be migrated further includes: Classifying the data to be processed according to a preset data classification program to obtain a second category of data to be processed; and Under the preset data scheduling platform, the preset data synchronization tool is used to process the second type of data to be processed to obtain the second data to be migrated.
6. The migration method according to claim 5, characterized in that: The preset data synchronization tool includes a preset data splitting program, and the data splitting program runs based on a preset data synchronization algorithm framework; The step of processing the second type of data to be processed in a preset data synchronization tool under the preset data scheduling platform to obtain the second data to be migrated includes: Based on the preset data scheduling platform and the preset data synchronization algorithm framework, the preset data splitting program is used to split the second type of data to be processed to obtain the second data to be migrated consisting of multiple groups of second sub-data to be migrated.
7. The migration method according to claim 6, characterized in that: The preset data migration tool further includes a preset distributed message queue, the preset second database includes a distributed data warehouse, and the using of the preset data migration tool to store the data to be migrated into the preset second database includes: Under the preset data scheduling platform, the preset data synchronization tool and the preset distributed message queue are used to store the second data to be migrated in the distributed data warehouse.
8. The migration method according to claim 6, characterized in that: The preset second database includes a distributed data warehouse and a distributed file system, and the migration method further includes: A preset data storage program is used to read the second data to be migrated in the distributed data warehouse, and the second data to be migrated is stored in the distributed file system.
9. The migration method according to claim 1, characterized in that: The data to be migrated includes third data to be migrated; and the using a preset data migration tool to process the data to be processed to obtain the data to be migrated further includes: Classifying the data to be processed according to a preset data classification program to obtain a third category of data to be processed; Under the preset data scheduling platform, receiving data processing instructions; and Under the preset data scheduling platform, the third type of data to be processed is processed according to the data processing instruction to obtain the third data to be migrated.
10. The migration method according to claim 9, characterized in that: The preset data migration tool includes a preset distributed message queue, the preset second database includes a column storage library, and using the preset data migration tool to store the data to be migrated into the preset second database includes: Under the preset data scheduling platform, receiving a first data migration instruction; and In response to the first data migration instruction, the third data to be migrated is stored in the column storage library using the preset distributed message queue.
11. The migration method according to claim 9, characterized in that: The preset second database includes a target disk, and using a preset data migration tool to store the data to be migrated into the preset second database includes: Under the preset data scheduling platform, receiving a second data migration instruction; and In response to the second data migration instruction, the third data to be migrated is stored in the target disk.
12. The migration method according to claim 9, characterized in that: The preset second database includes a distributed file system, and the using of a preset data migration tool to store the data to be migrated into the preset second database includes: Under the preset data scheduling platform, receiving a second data migration instruction; and According to the second data migration instruction, the third data to be migrated is stored in the distributed file system.
13. A big data migration system, characterized in that: The system comprises: A reading module, used for reading data in a preset first database to obtain data to be processed; A processing module, used to process the data to be processed using a preset data migration tool to obtain data to be migrated; and The data to be migrated is stored in a preset second database using a preset data migration tool.
14. The migration system according to claim 13, characterized in that: The data to be migrated includes first data to be migrated; the preset data migration tool includes a preset data import tool, and the processing module is further used for: Classifying the data to be processed according to a preset data classification program to obtain a first category of data to be processed; and Under the preset data scheduling platform, the first category of data to be processed is processed using a preset data processing program in the preset data import tool to obtain the first data to be migrated.
15. The migration system according to claim 14, characterized in that: The preset data processing program includes a preset data cleaning program, a preset data conversion program, and a preset data filtering program. The processing module is also used for: Under the preset data scheduling platform, using the preset data cleaning program, cleaning the first data to be processed to obtain data to be converted; Under the preset data scheduling platform, using the preset data conversion program, converting the data to be converted to obtain data to be filtered; and Under the preset data scheduling platform, the preset data filtering program is used to filter the first data to be filtered to obtain the first data to be migrated.
16. The migration system according to claim 14, characterized in that: The preset data migration tool further includes a preset distributed message queue, the preset second database includes a distributed file system, a distributed data warehouse, and a column storage library, and the processing module is further used for: Under the preset data scheduling platform, using the preset data import tool and the preset distributed message queue, the first data to be migrated is stored in the distributed file system; and According to a preset data distribution program, the preset data import tool and the preset distributed message queue are used to store the first data to be migrated stored in the distributed file system into the distributed data warehouse and the column storage library.
17. The migration system according to claim 13, characterized in that: The preset data migration tool includes a preset data synchronization tool, the data to be migrated includes second data to be migrated; and the processing module is further used for: Classifying the data to be processed according to a preset data classification program to obtain a second category of data to be processed; and Under the preset data scheduling platform, the preset data synchronization tool is used to process the second type of data to be processed to obtain the second data to be migrated.
18. The migration system according to claim 17, characterized in that: The preset data synchronization tool includes a preset data splitting program, and the data splitting program runs based on a preset data synchronization algorithm framework; the processing module is also used for: Based on the preset data scheduling platform and the preset data synchronization algorithm framework, the preset data splitting program is used to split the second type of data to be processed to obtain the second data to be migrated consisting of multiple groups of second sub-data to be migrated.
19. The migration system according to claim 18, characterized in that: The preset data migration tool further includes a preset distributed message queue, the preset second database includes a distributed data warehouse, and the processing module is further used for: Under the preset data scheduling platform, the preset data synchronization tool and the preset distributed message queue are used to store the second data to be migrated in the distributed data warehouse.
20. The migration system according to claim 18, characterized in that: The second database includes a distributed data warehouse and a distributed file system, and the processing module is further used for: A preset data storage program is used to read the second data to be migrated in the distributed data warehouse, and the second data to be migrated is stored in the distributed file system.
21. The migration system according to claim 13, characterized in that: The data to be migrated includes third data to be migrated; and the processing module is further used for: Classifying the data to be processed according to a preset data classification program to obtain a third category of data to be processed; Under the preset data scheduling platform, receive data processing instructions; and Under the preset data scheduling platform, the third type of data to be processed is processed according to the data processing instruction to obtain the third data to be migrated.
22. The migration system according to claim 21, characterized in that: The preset data migration tool includes a preset distributed message queue, the preset second database includes a column storage library, and the processing module is further used for: Under the preset data scheduling platform, receiving a first data migration instruction; and In response to the first data migration instruction, the third data to be migrated is stored in the column storage library using the preset distributed message queue.
23. The migration system according to claim 21, characterized in that: The preset second database includes a target disk, and the processing module is further used for: Under the preset data scheduling platform, receiving a second data migration instruction; and In response to the second data migration instruction, the third data to be migrated is stored in the target disk.
24. The migration system according to claim 21, characterized in that: The preset second database includes a distributed file system, and the processing module is further used for: Under the preset data scheduling platform, receiving a second data migration instruction; and According to the second data migration instruction, the third data to be migrated is stored in the distributed file system.
25. A chip, characterized in that: The chip includes a memory and a processor, the memory is configured to store a computer program, and the processor implements the migration method according to any one of claims 1 to 12 when executing the computer program.
26. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the migration method described in any one of claims 1 to 12 is implemented.