An Optimization Method and System for Cloud Data Migration
By slicing local private cloud resources by component type and mapping them into computing chains, combining Flink and data differentiated warning models, the problems of long and interruption of massive data on the cloud are solved, and efficient data migration and resource utilization are achieved.
Patent Information
- Application Number
- CN202211323617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-27
AI Technical Summary
In the prior art, it takes a long time to go to the cloud to use massive data, and the problem of data interruption due to exhaustion of basic local equipment resources or network abnormalities during the cloud is not effectively solved.
Slice local private cloud resources into independent private clouds according to component types, and synchronize resources through a multi-cloud management platform, create data pools and thread pools, map data cloud-based processes into computing chains, and use Flink to create TaskManager processes for priority thread allocation and data upload, and predict abnormal data based on data differentiated warning models.
It realizes the prediction of abnormal data and the handling of abnormal situations during the cloud-based process, reduces data interruptions and improves the efficiency and resource utilization of data on the cloud.
Smart Images

Figure CN115756308B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to an optimization method and system for cloud data migration. Background Art
[0002] In order to respond to the country's call for cloud adoption, and with the advent of the big data era, huge amounts of data fragments are generated every day in all walks of life. The data measurement units have developed from Byte, KB, MB, GB, TB to PB, EB, ZB, YB, and even BB, NB, DB for measurement. The collection of data in the big data era is no longer a technical problem. However, in the face of such a large amount of data, how can we find its internal laws?
[0003] The data lake architecture is for information storage facing multiple data sources, including the Internet of Things. Big data analysis or archiving can process data by accessing the data lake or deliver data subsets to requesting users. However, the data lake architecture is not just a huge disk. The data persistence and security of the data lake are factors that need to be considered first.
[0004] Enterprise cloud adoption refers to the process in which an enterprise applies information infrastructure, management, business, etc. based on the Internet and connects to social resources, shared services, and capabilities through the Internet and cloud computing means. When an enterprise transforms into a platform-based organization, it should also move towards "enterprise cloud adoption", that is, transform itself into a "cloud organization". The industry has generally recognized that the trend of "overall enterprise cloud adoption" is irreversible.
[0005] In the prior art, migrating a large amount of data to the cloud takes a long time, and there are problems such as the interruption of cloud migration data due to the exhaustion of local device basic resources or network anomalies during the cloud migration process. Summary of the Invention
[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide an optimization method and system for cloud data migration, which maps the entire cloud data migration process into an operation chain, can predict whether the data is abnormal during the cloud migration process, and solves the problems that migrating a large amount of data to the cloud takes a long time and the interruption of cloud migration data due to the exhaustion of local device basic resources or network anomalies during the cloud migration process.
[0007] According to one aspect of the present invention, the present invention provides an optimization method for cloud data migration, and the method includes the following steps:
[0008] S1: Segment and manage local private cloud resources according to component types to form independent private clouds, and form a multi-cloud management platform by combining all the independent private clouds;
[0009] S2: Create a data pool, and put the physical device information, network data, application operation data, and log text data associated with the private cloud of the multi-cloud management platform into the initial data pool through an open interface;
[0010] S3: Create a thread pool, map the data values in each type of data pool to operators, and map the entire process of data cloudification to an operation chain.
[0011] Preferably, the step S1 further includes:
[0012] Slice and manage all local private cloud resources through fusioncomputer, and complete the resource synchronization between the fusioncomputer and the multi-cloud management platform through an open interface.
[0013] Preferably, the step S1 further includes: constructing a data differentiation early warning model, and predicting the probability of abnormal data occurrence and the probability of normal data occurrence through the data differentiation early warning model; the data differentiation early warning model is:
[0014] X(k + 1) = X(k) × P
[0015] Wherein, X(k) represents the state vector of the trend analysis and prediction object at the k-th moment, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at the (k + 1)-th moment.
[0016] Preferably, the data pool includes an analog signal data pool, an application program data pool, and a text data pool; the data in the data pool includes metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pool is: numerical value metadata ID meta-process data ID virtualization identification ID from multiple platforms failure probability value; wherein, is a numerical separator.
[0017] Preferably, the creating of the thread pool, mapping the data values in each type of data pool to operators, and mapping the entire process of data cloudification to an operation chain includes:
[0018] Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads according to the failure probability value;
[0019] Send the operation chain threads with the failure probability values of the operators in the operation chain less than the preset threshold to multiple shared task time slots in sequence and evenly;
[0020] Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
[0021] According to another aspect of the present invention, the present invention further provides an optimization system for cloud data migration, and the system includes:
[0022] A privatization management module, which is used to slice and manage local private cloud resources according to component types, form independent private clouds, and form a multi-cloud management platform by combining all the independent private clouds;
[0023] A storage classification module, which is used to create a data pool, and put physical device information, network data, application operation data, and log text data associated with the private clouds of the multi-cloud management platform into the initial data pool through an open interface;
[0024] A cloud acceleration module, which is used to create a thread pool, map the data values in each type of data pool to operators, and map the entire data cloudification process to an operation chain.
[0025] Preferably, the privatization management module is further used for:
[0026] Slice and manage all local private cloud resources through fusioncomputer, and complete resource synchronization between the fusioncomputer and the multi-cloud management platform through an open interface.
[0027] Preferably, the privatization management module is further used for: constructing a data difference warning model, and predicting the probability of abnormal data occurrence and the probability of normal data occurrence through the data difference warning model; the data difference warning model is:
[0028] X(k + 1) = X(k) × P
[0029] Wherein, X(k) represents the state vector of the trend analysis and prediction object at the k-th moment, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at the (k + 1)-th moment.
[0030] Preferably, the data pools created by the storage classification module include an analog signal data pool, an application program data pool, and a text data pool; the data in the data pools include metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pools is: numerical value metadata ID meta-process data ID virtualization identification ID from multiple platforms failure probability value; wherein, is a numerical separator.
[0031] Preferably, the cloud acceleration module creates a thread pool, and mapping the data values in each type of data pool to operators and mapping the entire data cloudification process to an operation chain includes:
[0032] Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads based on the failure probability value;
[0033] For the operation chain threads with the failure probability values of each operator in the operation chain less than the preset threshold, send them evenly to multiple shared task time slots in sequence;
[0034] Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
[0035] Beneficial effects: The present invention classifies data by type using the data lake technology, combines the characteristics of accelerating data operator chains in Flink, maps the entire data cloud - uploading process to an operation chain, can predict whether the data is abnormal during the cloud - uploading process, and solves the problems that the cloud - uploading of massive data takes a long time and the cloud - uploading data is interrupted due to the exhaustion of local device basic resources or network anomalies during the cloud - uploading process.
[0036] The features and advantages of the present invention will become clear by referring to the following drawings and the detailed description of the specific embodiments of the present invention. Description of the Drawings
[0037] Figure 1 It is a flowchart of the optimization method for cloud data migration;
[0038] Figure 2 It is a schematic diagram of the optimization system for cloud data migration. Detailed Embodiments
[0039] Next, in combination with the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] Figure 1 It is a flowchart of the optimization method for cloud data migration. As Figure 1 shown, this embodiment provides an optimization method for cloud data migration, and the method includes the following steps:
[0042] S1: Cut and manage the local private cloud resources according to component types to form independent private clouds, and form a multi - cloud management platform with all the independent private clouds.
[0043] Preferably, the step S1 further includes:
[0044] The FusionCompute is used to slice and manage all local private cloud resources, and resource synchronization between the FusionCompute and the multi-cloud management platform is completed through open interfaces.
[0045] Specifically, all local private cloud resources are sliced and managed by the FusionCompute technology according to different types (components such as VMware, KVM, OpenStack, etc.) to form independent private clouds, and all private clouds are combined to form a multi-cloud management platform. And resource synchronization between the FusionCompute and the multi-cloud management platform is completed through the OpenAPI.
[0046] Among them, FusionCompute is a virtualization product of Huawei, which is developed with corresponding feature enhancements based on open source technologies.
[0047] In this step, first, slicing of all local private cloud resources of the FusionCompute is completed, and resource synchronization between the FusionCompute and the multi-cloud management platform is completed through the OpenAPI. Secondly, while data interaction occurs between the OpenAPI and the private cloud resources of the multi-cloud management platform, the problem of exposure of private data of independent private clouds is also protected accordingly.
[0048] Preferably, the step S1 further includes: constructing a data difference early warning model, and predicting the probability of abnormal data occurrence and the probability of normal data occurrence through the data difference early warning model; the data difference early warning model is:
[0049] X(k + 1) = X(k) × P
[0050] Among them, X(k) represents the state vector of the trend analysis and prediction object at the k-th moment, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at the (k + 1)-th moment.
[0051] Specifically, for example, assume that the total amount of historical data transmission is 1, among which the proportion of abnormal data is 30% (0.3), and the proportion of normal data is 70% (0.7). Also, according to historical data analysis, it is obtained that 60% of the current abnormal data may continue to be abnormal, and 40% may turn into normal.
[0052] And among the current normal data with a proportion of 70%, 70% may still be normal, while 30% will become abnormal. The calculation process is as follows:
[0053] The probability of abnormal data occurrence next time: 0.3 × 0.6 + 0.3 × 0.7 = 0.39 × 100 = 39%;
[0054] Next normal data occurrence probability: 0.3 x 0.4 + 0.7 x 0.7 = 0.61 x 100 = 61%.
[0055] S2: Create a data pool, and put the physical device information, network data, application operation data, and log text data associated with the private cloud of the multi-cloud management platform into the initial data pool through an open interface.
[0056] Preferably, the data pool includes an analog signal data pool, an application data pool, and a text data pool; the data in the data pool includes metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pool is: numerical value metadata ID meta-process data ID virtualization identification ID from multiple platforms failure probability value; where is a numerical separator.
[0057] Specifically, deploy the initial data pool on the server, and create three major types of data pools to obtain the data sorted from the initial data pool.
[0058] The three major types of data pools are the analog signal data pool, the application data pool, and the text data pool, and the stored data is trimmed.
[0059] Put the physical devices and network data, application operation data, and log text data associated with the private cloud of the multi-cloud management platform into the initial data pool through openapi, and at the same time capture the metadata corresponding to the collected data. The purpose of setting up the initial data pool is to act as a data storage unit and prepare for the next step of data entering different types of data pools according to data characteristics. Map the metadata, meta-process data, and the three-party relationship between the data and the associated metadata and meta-process data to a metadata identifier and pass it to the corresponding type of data processing pool together.
[0060] Virtual identification format: numerical value metadata ID meta-process data ID specific virtualization identification ID from multiple platforms failure probability value (kvm\vmware\openstack). Where
[0061] : is a numerical separator.
[0062] Metadata: Descriptions of data records, indexes, key values, and relationships between different data attributes, etc.
[0063] Meta-process data: More valuable for analysis than the collected data, usually containing richer information, including records, dates, locations, responsible persons, recording devices, and other attached information.
[0064] S3: Create a thread pool, map the data values in each type of data pool to operators, and map the entire process of uploading data to the cloud to an operation chain.
[0065] Preferably, creating a thread pool, mapping the data values in each type of data pool to operators, and mapping the entire data cloud - uploading process to an operation chain includes:
[0066] Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads based on the failure probability value;
[0067] Send the operation chain threads with the failure probability values of each operator in the operation chain less than the preset threshold to multiple shared task time slots evenly in sequence;
[0068] Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
[0069] Specifically, create a thread pool and map the data values (metadata, meta - process data, virtualization identification ID, failure probability value) of each type of data pool in step S2 to operators, and map the entire cloud - uploading process to an operation chain.
[0070] First, each operation chain is mapped to a TaskManager process. The JobManager manages the private - cloud data operations (TaskManager) of the multi - cloud management platform through the Actor System, and allocates priority threads according to the failure probability value, so as to fully utilize the resources of Task Slots (process slots) so that larger sub - tasks can be evenly distributed on the TaskManager, and the Client is used to manage the data flow to the JobManagers.
[0071] Second, preferentially send the operation chain threads with the failure probability values of each operator in the operation chain (SubTask) not greater than 50% to multiple shared Task Slots evenly in sequence. Thus, the resources of Task Slots (process slots) can be fully utilized so that larger sub - tasks can be evenly distributed on the TaskManager, and the TCP connections and heartbeat messages of other TaskManagers under the same JobManager can be shared. At the same time, some data sets and data structures can be shared, thereby reducing the task overhead.
[0072] Then, upload the data sorted out by the process pool to the corresponding public cloud according to the passed - in virtualization identification ID parameter.
[0073] For example: Different public clouds correspond to the importance of the services to which the local private - cloud data belongs. The private - cloud VMware data is correspondingly uploaded to the Huawei Cloud in the public cloud.
[0074] The thread of this embodiment is a single sequential control flow in a process. Multiple threads can be concurrent in a process, and each thread executes different tasks in parallel. In an application, threads need to be used multiple times, which means that threads need to be created and destroyed multiple times, and the creation and destruction of threads will inevitably consume memory. The thread pool can reuse threads, avoid repeated creation and destruction of threads as much as possible, and reduce memory consumption.
[0075] This embodiment uses data lake technology to classify data by type, and combines the data operator chain acceleration feature of Flink to map the entire data cloud migration process into an operation chain. It can predict whether the data is abnormal during the cloud migration process, and solves the problem that it takes a long time to migrate massive data to the cloud, and the cloud migration data is interrupted due to the exhaustion of local device basic resources or network abnormalities during the cloud migration process.
[0076] Flink is a framework and distributed processing engine for stateful computations on unbounded and bounded data streams. Flink is designed to run in all common cluster environments, performing computations at memory speed and at any scale. Flink's TaskManager is responsible for mapping each computation chain to a TaskManager process; JobManager manages the data computations in distributed disaster recovery drills through Actor System and participates in the central early warning model computation to obtain the failure probability value of each operator, and allocates priority threads based on the probability value, thereby making full use of TaskSlots resources so that larger operators can be evenly distributed on TaskManagers; finally, Client is used to manage the flow of local disaster recovery drill data to JobManagers.
[0077] Example 2
[0078] Figure 2 This is a schematic diagram of the optimized system for cloud data migration. Figure 2 As shown, this embodiment also provides an optimization system for cloud data migration, the system comprising:
[0079] The privatization management module 201 is used to divide and manage local private cloud resources according to component types to form independent private clouds, and to form a multi-cloud management platform with all the independent private clouds;
[0080] The storage classification module 202 is used to create a data pool, and put the physical device information and network data, application operation data, and log text data associated with the private cloud of the multi-cloud management platform into the initial data pool through an open interface;
[0081] The cloud migration acceleration module 203 is used to create a thread pool, map the data values in each type of data pool into operators, and map the entire data migration process into a calculation chain.
[0082] Preferably, the privatization management module 201 is further configured to:
[0083] Slice and manage all local private cloud resources through fusioncomputer, and complete resource synchronization between the fusioncomputer and the multi-cloud management platform through an open interface.
[0084] Preferably, the privatization management module 201 is further configured to: construct a data difference early warning model, and predict the probability of abnormal data occurrence and the probability of normal data occurrence through the data difference early warning model; the data difference early warning model is:
[0085] X(k + 1) = X(k) × P
[0086] where X(k) represents the state vector of the trend analysis and prediction object at time k, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at time k + 1.
[0087] Preferably, the data pool created by the storage classification module 202 includes an analog signal data pool, an application data pool, and a text data pool; the data in the data pool includes metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pool is: numerical value metadata ID meta-process data ID virtualization identification ID from multiple platforms failure probability value.
[0088] Preferably, the cloud acceleration module 203 creates a thread pool, maps the data values in each type of data pool to operators, and maps the entire data cloud process to an operation chain, including:
[0089] Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads according to the failure probability value;
[0090] Send the operation chain threads with the failure probability values of each operator in the operation chain less than the preset threshold to multiple shared task time slots evenly in sequence;
[0091] Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
[0092] The specific implementation processes of the functions implemented by each module in this Embodiment 2 are the same as those in Embodiment 1, and will not be elaborated here.
[0093] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. An optimization method for cloud data migration, characterized in that, The method includes the following steps: S1: Split and manage the local private cloud resources according to component types to form independent private clouds, and form a multi-cloud management platform by combining all the independent private clouds; S2: Create a data pool, and put the physical device information, network data, application operation data, and log text data associated with the private clouds of the multi-cloud management platform into the initial data pool through an open interface; S3: Create a thread pool, map the data values in each type of data pool to operators, and map the entire process of uploading data to the cloud to an operation chain; Among them, step S1 further includes: splitting and managing all local private cloud resources through fusioncomputer, and completing resource synchronization between the fusioncomputer and the multi-cloud management platform through an open interface; Among them, step S1 further includes: constructing a data difference warning model, and predicting the probability of abnormal data occurrence and the probability of normal data occurrence through the data difference warning model; the data difference warning model is: X(k + 1)=X(k)×P Where X(k) represents the state vector of the trend analysis and prediction object at time k, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at time k + 1; Among them, the creation of the thread pool, mapping the data values in each type of data pool to operators, and mapping the entire process of uploading data to the cloud to an operation chain includes: Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads according to the failure probability value; Send the operation chain threads with the failure probability values of each operator in the operation chain less than the preset threshold to multiple shared task time slots in sequence and evenly; Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
2. The method according to claim 1, wherein The data pool includes an analog signal data pool, an application data pool, and a text data pool; the data in the data pool includes metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pool is: numerical valuemetadata IDmeta-process data IDvirtualization identification ID from multiple platformsfailure probability value; where is a numerical separator.
3. An optimized system for cloud data migration, characterized in that, The system includes: A privatization management module, which is used to split and manage the local private cloud resources according to component types to form independent private clouds, and form a multi-cloud management platform by combining all the independent private clouds; A storage classification module, which is used to create a data pool, and put the physical device information, network data, application operation data, and log text data associated with the private clouds of the multi-cloud management platform into the initial data pool through an open interface; A cloud upload acceleration module, which is used to create a thread pool, map the data values in each type of data pool to operators, and map the entire process of uploading data to the cloud to an operation chain; Among them, the privatization management module is further configured to: slice and manage all local private cloud resources through fusioncomputer, and complete resource synchronization between the fusioncomputer and the multi-cloud management platform through an open interface; Among them, the privatization management module is further configured to: construct a data difference warning model, and predict the probability of abnormal data occurrence and the probability of normal data occurrence through the data difference warning model; the data difference warning model is: X(k + 1) = X(k) × P where X(k) represents the state vector of the trend analysis and prediction object at time k, P represents the one-step transition probability matrix, and X(k + 1) represents the state vector of the trend analysis and prediction object at time k + 1; Among them, the cloud acceleration module creates a thread pool, maps the data values in each type of data pool to operators, and maps the entire data cloud process to an operation chain, including: Create a thread pool through Flink, map each operation chain to a TaskManager process, and allocate priority threads according to the failure probability value; Uniformly send the operation chain threads with the failure probability values of each operator in the operation chain less than the preset threshold to multiple shared task time slots in sequence; Upload the data processed by the process pool to the corresponding public cloud according to the virtualization identification ID.
4. The system according to claim 3, characterized in that The data pools created by the storage classification module include an analog signal data pool, an application data pool, and a text data pool; the data in the data pools include metadata, meta-process data, virtualization identification ID, and failure probability value; the format of the data in the data pools is: numerical value metadata ID meta-process data ID virtualization identification ID from multiple platforms failure probability value; where is the numerical separator.
Citation Information
Patent Citations
Cloud and mist mixing path determination method for medical big data
CN110830292A
K8s network fault prediction method based on Markov chain and Bayesian network
CN115037634A