A Big Data Migration Optimization Method Combining Distributed Clusters and Virtual Machine Rooms
By creating a virtual data center, building a weight-first model and marking sorting during the big data migration process, combining the technology of data pools and storage tanks, the data differentiation problem in big data migration is solved, and efficient and orderly data migration is achieved.
Patent Information
- Application Number
- CN202211361207.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-11-02
AI Technical Summary
The existing technology failed to effectively solve the problem of data differentiation during the migration of big data to other virtual machine rooms, resulting in low efficiency in synchronous data and cumbersome configuration and incremental data operations.
By creating a virtual data center and combining cluster management, Pareto optimal deconstructs weight-first model and mark-copying, and finally using data pools and storage tanks to combine meta-process data for classified migration.
It realizes orderly and efficient migration of data in the order of marking priorities during the big data migration process, fixes the differentiated problems arising from data migration, and improves migration efficiency.
Smart Images

Figure CN115509461B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data and AI, and particularly relates to a method for optimizing big data migration by combining a distributed cluster with a virtual computer room. Background Art
[0002] With the development of science and technology and the Internet, it has promoted the advent of the big data era. Vast amounts of data fragments are generated every day in all walks of life, and the data measurement units have developed from Byte, KB, MB, GB, TB to PB, EB, ZB, YB, and even BB, NB, DB for measurement. The acquisition of data in the big data era is no longer a technical problem. However, in the face of such a large amount of data, how can we find its internal laws? FusionComputer is an enterprise-level open server virtualization solution verified in the cloud computing environment, which can transform a static and complex IT environment into a more dynamic and manageable virtual data center. At the same time, it can also provide advanced management functions to achieve the integration and automation of the virtual data center, while the cost is much lower than other solutions.
[0003] Chinese invention patent CN202110585631.4 discloses a data migration method, device, storage medium, and data migration device. The data migration method is applied to a data migration system, which includes a source cluster, a target cluster, and a message queue. The data migration method includes: obtaining a migration instruction, and synchronizing the stock data from the source cluster to the target cluster according to the migration instruction, and writing the incremental data to the source cluster and the target cluster; when an exception occurs in writing the incremental data to the target cluster, determining the data with an exception in writing to the target cluster in the incremental data as target data, and writing the target data to the message queue; obtaining a merge instruction, and merging the target data into the target cluster according to the merge instruction and the data version of the target data, thereby improving the efficiency of data migration and reducing the participation of users in the data migration process.
[0004] Chinese invention patent CN202010093662.3 discloses a data migration method, including: establishing a connection between a source object storage cluster and a target object storage cluster through an S3 protocol interface; using the S3 protocol interface to read a user list or a bucket list from the source object storage cluster; allocating data migration tasks according to the user list or the bucket list; and executing the data migration tasks, reading object data and writing it to the target object storage cluster in real time. This application simplifies the data migration operation and greatly improves the data migration efficiency.
[0005] However, the prior art does not solve the problem of data differentiation generated during the migration of big data to other virtual computer rooms. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides an optimization method for big data migration combining a distributed cluster and a virtual computer room. The present invention mainly combines the creation of a virtual data center with cluster management, then constructs a weight priority model and label sorting using the Pareto optimal solution, and finally classifies and migrates data using a data pool and a storage tank in combination with meta-process data, which can combine artificial intelligence in the field of big data migration to solve the differential problems generated during the data migration process, and migrate the migrated data in an orderly and efficient manner.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An optimization method for big data migration combining a distributed cluster and a virtual computer room, comprising the following steps:
[0009] S11. Create a virtual data center combined with a cluster management module;
[0010] S12. Construct a weight priority model and label sorting using the Pareto optimal solution;
[0011] S13. Classify and migrate data using a data pool and a storage tank in combination with meta-process data.
[0012] The virtual data center can abstract and integrate physical resources through virtualization technology, dynamically allocate and schedule resources, realize the automated deployment of the data center, and will greatly reduce the operating cost of the data center;
[0013] The weight priority model constructed by the Pareto optimal solution is used to analyze all Node data of Pod business transaction data with a large volume and related to the core application IP, and obtain the optimal solution Node; the label sorting is used for the priority labeling of Nodes (nodes), and ID is the priority label ( represents a number, the priority of 0 is the highest, and the higher the number, the lower the priority);
[0014] The data pool is a structure for storing data, used for pre-processing, caching, and reference checking of a large amount of data. It is different from business process data and the data middle platform for analysis. It is an auxiliary database of the business processing system.
[0015] For a further description of the present invention, the step S12 includes the following steps:
[0016] S121. Construct a weight priority model using the Pareto analysis method (Pareto);
[0017] S122. After analyzing all Node data of Pod business transaction data with a large volume and related to the core application IP through the weight priority model to obtain the optimal solution Node, perform priority labeling on the Node;
[0018] S123. Independently migrate the optimal solution Node data through the optimal memory pool opened by Flink.
[0019] The main purpose of the weight - priority model is to better label the Nodes (nodes) with mixed services during the migration process without affecting the migration of other Nodes (nodes), so as to migrate the migration data in an orderly and efficient manner according to the label - priority order.
[0020] The weight of the volume of business transaction data is obtained by analyzing the log files of the customer's core business data to find out which associated Nodes (nodes) it often communicates with.
[0021] The attributes of Flink include: JobManager: responsible for coordinating local privacy and public data and transmitting them to the central database in the form of distributed tasks through a thread pool; TaskManager: responsible for mapping to a process for local - to - central data transmission, storing Task Slots, and dividing the memory into multiple Task Slots (process memory slots) and connecting them to memory resources; Task Slots (process memory slots) are spaces allocated from memory to store processes; Flink is used to streamline the data flow to facilitate the orderly and efficient migration of migration data according to the label - priority order.
[0022] For further description of the present invention, step S13 includes the following steps:
[0023] S131. Deploy a central initial data pool on the central server.
[0024] S132. Bind initial data storage tanks to each Node in the distributed virtual computer room to collect Pod data managed by each local Node, and conduct preliminary sorting. Put the data with little value into the miscellaneous data pool allocated from the central initial data pool.
[0025] S133. Through the priority label in step S12, perform weight - priority model operations on the Node with the highest priority and all Pods it manages according to the log data stored in the Node, and classify them according to the conditions of data size and the number of times of associated core application IPs to obtain the Pareto optimal solution and update the priority label.
[0026] S134. Migrate the migration data according to the label - priority order.
[0027] The central initial data pool is used to pre - process the initial data for caching for future reference.
[0028] The data with little value includes data with little fluctuation and a large amount of repetition, which is judged to have little value from the perspective of value analysis, such as the data collected during normal monitoring; after being stored in the miscellaneous data pool for a certain period of time, it is finally put into a storage device for easy access when querying historical data later, without migration.
[0029] The Pareto optimal solution refers to a result where it is impossible to reallocate resources to make at least one person's situation better without making the situation of anyone else worse.
[0030] For a further description of the present invention, the central initial data pool is composed of multiple data storage tanks, and each data storage tank corresponds to a cluster of Nodes.
[0031] For a further description of the present invention, the weight priority model formula is:
[0032] minf(x)=[f 1 (x),…,f p (x)]
[0033] In the formula: f p (x) represents the p-th objective function, x represents an n-dimensional decision vector, and x=[x 1 ,x 2 …,x n .
[0034] For a further description of the present invention, the step S11 includes the following steps:
[0035] S111. Create a Kubernetes central cluster on the central server;
[0036] S112. Use FusionComputer technology to partition the physical resources of the central server computer room, create multiple virtual computer rooms, and map each virtual computer room to a unique virtual computer room cluster identifier with the Node.
[0037] Kubernetes is an open-source Linux container automated operation and maintenance platform for managing a cluster composed of multiple hosts;
[0038] FusionComputer is an enterprise-level open server virtualization solution that has been verified in a cloud computing environment. It can transform a static and complex IT environment into a more dynamic and manageable virtual data center, thus greatly reducing the cost of the data center. At the same time, it can provide advanced management functions to achieve the integration and automation of the virtual data center, and the cost is much lower than other solutions.
[0039] Further description of the present invention, the Kubernetes-based cluster includes a Master, Nodes, and Pods;
[0040] The Master is used to manage local Nodes;
[0041] The Nodes are used to execute the assigned Pods under the control of the Kubernetes master node;
[0042] The Pods are used to run the data generated by business applications.
[0043] Further description of the present invention, the meta-process data includes recorded date, location, responsible person, equipment, and other accessory information.
[0044] The present invention has the following beneficial effects:
[0045] 1. The present invention uses the Pareto optimal solution to construct a weight priority model and label sorting, realizing more optimal labeling of Nodes (nodes) with mixed services during the migration process without affecting the migration of other Nodes (nodes), and then streamlining the data process through Flink technology, so as to migrate the migration data in an orderly and efficient manner according to the label priority order.
[0046] 2. The present invention uses a data pool and storage tanks to classify and migrate data in combination with meta-process data, fixing the data differences generated during the data migration process and the problems of low synchronization data efficiency, cumbersome configuration, and incremental data operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flowchart of an optimization method for big data migration of a distributed cluster combined with a virtual computer room.
[0048] Figure 2 It is a model diagram of an optimization method for big data migration of a distributed cluster combined with a virtual computer room.
[0049] Figure 3 It is Figure 1 a flowchart of creating a virtual data center combined with a cluster management module in
[0050] Figure 4 It is Figure 1 a flowchart of constructing a weight priority model and label sorting using the Pareto optimal solution in
[0051] Figure 5 It is Figure 1 a flowchart of classifying and migrating data using a data pool and storage tanks in combination with meta-process data in DETAILED DESCRIPTION OF THE INVENTION
[0052] The present invention will be further described below with reference to the accompanying drawings.
[0053] A distributed cluster combined with a virtual computer room big data migration optimization method, the process of which is as follows Figure 1 shown, and its model is as follows Figure 2 shown, including the following steps:
[0054] S11. Create a virtual data center combined with a cluster management module; the specific steps are as follows Figure 3 shown:
[0055] S111. Create a Kubernetes central cluster on the central server, which mainly includes three objects: Master, Node, and Pod; the Master (main node) is used to manage each local Node; the Node (node) is used to execute the assigned Pod under the control of the Kubernetes master node; the Pod is used to run business applications to generate a large amount of data. During the data migration process, mainly the Pod is migrated to other Nodes (nodes);
[0056] S112. Use the FusionComputer technology to partition the physical resources of the central server computer room, create multiple virtual computer rooms, and map each virtual computer room to a unique virtual computer room cluster identifier with the Node.
[0057] S12. Use the Pareto optimal solution to construct a weight priority model and label sorting; the specific steps are as follows Figure 4 shown:
[0058] S121. Use the Pareto analysis method (Pareto) to construct a weight priority model; the main purpose of using the model is to perform better labeling on the Nodes (nodes) with mixed services during the migration process without affecting the migration of other Nodes (nodes), so as to migrate the migration data in an orderly and efficient manner according to the label priority order;
[0059] S122. Analyze all the Node data of the Pod business with a large volume of transactions and related to the core application IP through the weight priority model to obtain the optimal solution Node, and then perform priority labeling on the Node; the priority label is marked as Node (node) ID (where is a number, 0 has the highest priority, and the higher the number, the lower the priority); the weight of the business transaction volume is obtained by analyzing the log files of the customer's core business data to find out which associated Nodes (nodes) are often involved.
[0060] S123. Independently migrate the optimal solution Node data through the optimal memory pool opened by Flink.
[0061] The formula of the weight priority model is:
[0062] min f(x) = [f 1 (x), …, f p (x)]
[0063] where: f p (x) represents the p-th objective function, x represents an n-dimensional decision vector, x = [x 1 , x 2 …, x n ;
[0064] For the multi-objective programming optimal solution problem to be solved by the weight priority model, denote its variable feasible region as S, and the corresponding objective feasible region Z = f(S).
[0065] Given a feasible point x * ∈ S, there is If f(x * ) < f(x), then x * is called the absolute optimal solution of the multi-objective programming problem. If there does not exist x ∈ S such that f(x * ) < f(x), then x * is called the effective solution of the objective programming problem. The effective solution of the multi-objective programming problem is also called the Pareto optimal solution.
[0066] S13. Classify and migrate data using a data pool and storage tanks in combination with meta-process data; the meta-process data includes the recorded date, location, responsible person, equipment, and other accessory information.
[0067] The data differentiation during the virtual machine room migration process is mainly reflected in the low efficiency of synchronizing data in the existing migration scheme and the cumbersome configuration and incremental data operations. Therefore, classify and migrate data using a data pool and storage tanks in combination with the original process data, thus fixing the differentiation problem generated during the data migration process;
[0068] As Figure 5 shown, step S13 specifically includes the following steps:
[0069] S131. Deploy a central initial data pool on the central server; the central initial data pool consists of multiple data storage tanks, and each data storage tank corresponds to a Node in a cluster;
[0070] S132. Deploy initial data storage tanks on each Node bound to the distributed virtual machine room to collect Pod data managed by each local Node, and perform preliminary sorting. Put the data with little value into the miscellaneous data pool allocated from the central initial data pool; after a certain period of storage, the miscellaneous data pool is finally put into the storage device for easy access when querying historical data later and is not migrated.
[0071] S133. Through the priority marking in step S12, perform a weighted priority model operation on the Node with the highest priority and all the Pods it manages according to the log data stored in the Node, and classify them according to the conditions of data size and the number of times associated with the core application IP to obtain Pareto optimal solutions (multiple optimal solutions may be obtained), and update the priority marking;
[0072] S134. Migrate the migration data in the order of marking priority. The Node and storage tank where the data with high weight is stored are migrated first. Before migration, use Huawei's FusionComputer technology to split the physical device into several virtual computer rooms and map and bind them to the Node.
[0073] The metadata is a description of the records, indexes, key values of the data, and the relationships between different data attributes, etc.
[0074] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.
Claims
1. A method for optimizing big data migration by combining a distributed cluster with a virtual computer room, characterized in that it includes the following steps: S11. Create a virtual data center combined with a cluster management module; S12. Use the Pareto optimal solution to construct a weight priority model and tag sorting; S13. Use a data pool and a storage tank to classify and migrate data by combining meta-process data; The step S12 includes the following steps: S121. Use the Pareto analysis method (Pareto) to construct a weight priority model; S122. Analyze all Node data of Pod business transaction data with a large volume and related to the core application IP through the weight priority model to obtain the optimal solution Node, and then preferentially tag the Node; S123. Independently migrate the optimal solution Node data through the optimal memory pool opened by Flink; The step S13 includes the following steps: S131. Deploy a central initial data pool on the central server; S132. Deploy an initial data storage tank on each Node bound to the distributed virtual computer room to collect Pod data managed by each local Node, and conduct preliminary sorting. Put the data with little value into the miscellaneous data pool allocated from the central initial data pool; S133. Through the priority tag in step S12, perform weight priority model operations on the highest-priority Node and all Pods it manages according to the log data stored in the Node, and classify them according to the conditions of data size and the number of times related to the core application IP to obtain the Pareto optimal solution and update the priority tag; S134. Migrate the migration data in the order of tag priority; The formula of the weight priority model is: ; where: f p (x) represents the p-th objective function, x represents an n-dimensional decision vector, x = [x 1 , x 2 …, x n .
2. The method for optimizing big data migration by combining a distributed cluster with a virtual computer room according to claim 1, characterized in that: The central initial data pool is composed of multiple data storage tanks, and each data storage tank corresponds to a Node of a cluster.
3. The method for optimizing big data migration by combining a distributed cluster with a virtual computer room according to claim 1, characterized in that: The step S11 includes the following steps: S111. Create a Kubernetes central cluster on the central server; S112. Use the FusionComputer technology to partition the physical resources of the central server computer room to create multiple virtual computer rooms, and map each virtual computer room to a unique virtual computer room cluster identifier with the Node.
4. The method for optimizing big data migration by combining a distributed cluster with a virtual computer room according to claim 3, characterized in that: A cluster based on Kubernetes includes a Master, a Node, and a Pod; The Master is used to manage each local Node; The Node is used to execute the assigned Pod under the control of the Kubernetes master node; The Pod is used to run business applications to generate data.
5. The method for optimizing big data migration by combining a distributed cluster with a virtual computer room according to claim 4, characterized in that: The meta-process data includes the recorded date, location, responsible person, and equipment.
Citation Information
Patent Citations
Data migration method and system, storage medium and data migration terminal
CN111309259A
Data migration method and device, storage medium and data migration equipment
CN113438275A
Optimized deploying method based on virtual cluster online migration
CN106020934A
Distributed storage multi-site synchronization optimization method and device, equipment and storage medium
CN114153397A