Data migration method and device, electronic equipment, storage medium and program product

By obtaining metadata and environment variables, determining entity types and dependencies, and generating and deploying data packets, the rapid, accurate and complete migration of data warehouses is achieved, and the problem of difficult to guarantee the migration time and accuracy in the existing technology is solved.

CN120492423APending Publication Date: 2025-08-15HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410176617.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-15

Smart Images

  • Figure CN120492423A_ABST
    Figure CN120492423A_ABST
Patent Text Reader

Abstract

The invention provides a data migration method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of cloud computing, the method comprises the following steps: firstly, obtaining metadata of a to-be-migrated object from a source data warehouse; then, according to the metadata, determining an entity type of the to-be-migrated object and a deployment dependency relationship between entities; acquiring an environment variable of the to-be-migrated object in the source data warehouse; and finally, migrating the to-be-migrated object to the target data warehouse based on the entity type, the deployment dependency relationship between the entities and the environment variable. In the embodiment of the invention, the entity type of the to-be-migrated object, the deployment dependency relationship between the entities and the environment variable are obtained from the source data warehouse, and the to-be-migrated object is migrated to the target data warehouse based on the entity type, the deployment dependency relationship between the entities and the environment variable, so that rapid migration of data can be realized; the efficiency and quality of data migration are improved, and meanwhile the accuracy and integrity of data are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to a data migration method, device, electronic device, storage medium, and program product. Background Art

[0002] A data warehouse is a computer system used to store, manage, and analyze data. Data warehouse migration refers to the process of migrating a data warehouse from one environment or platform to another. Time constraints and the diversity of the migration targets (such as batch scheduling development platforms, real-time computing development platforms, and various databases) can pose significant challenges when migrating data warehouses.

[0003] The development and implementation of business data warehouses often involve multiple cloud products. In the current business data warehouse development and implementation process, the migration of all or part of the data warehouse requires manual operation, which is time-consuming and prone to data loss or corruption, making it difficult to guarantee accuracy and integrity. Summary of the Invention

[0004] Embodiments of the present application provide a data migration method, apparatus, electronic device, storage medium, and program product to address the problem that manual data migration is time-consuming and difficult to ensure accuracy and completeness.

[0005] In a first aspect, an embodiment of the present application provides a data migration method, which includes: obtaining metadata of the object to be migrated from a source data warehouse; determining the entity type of the object to be migrated and the deployment dependency relationship between entities based on the metadata; obtaining the environment variables of the object to be migrated in the source data warehouse; and migrating the object to be migrated to a target data warehouse based on the entity type, the deployment dependency relationship between entities and the environment variables.

[0006] In the second aspect, an embodiment of the present application provides a data migration device, which includes: a first acquisition module for acquiring metadata of the object to be migrated from the source data warehouse; a determination module for determining the entity type of the object to be migrated and the deployment dependency relationship between entities based on the metadata; a second acquisition module for acquiring the environment variables of the object to be migrated in the source data warehouse; and a migration module for migrating the object to be migrated to the target data warehouse based on the entity type, the deployment dependency relationship between entities and the environment variables.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the above methods when executing the computer program.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.

[0010] Compared with the prior art, this application has the following advantages:

[0011] The present application provides a data migration method, apparatus, electronic device, storage medium, and program product. First, metadata of an object to be migrated is obtained from a source data warehouse. Then, based on the metadata, the entity type of the object to be migrated and the deployment dependency relationship between entities are determined. The environment variables of the object to be migrated in the source data warehouse are obtained. Finally, based on the entity type, the deployment dependency relationship between entities, and the environment variables, the object to be migrated is migrated to a target data warehouse. In an embodiment of the present application, the entity type of the object to be migrated, the deployment dependency relationship between entities, and the environment variables are obtained from the source data warehouse, and the object to be migrated is migrated to the target data warehouse based on the entity type, the deployment dependency relationship between entities, and the environment variables. This can achieve rapid data migration, improve the efficiency and quality of data migration, and ensure the accuracy and integrity of the data.

[0012] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0014] Figure 1 A schematic diagram of an application scenario of the data migration method provided in this application.

[0015] Figure 2 A schematic diagram of entities and dependency information according to an embodiment of the present application.

[0016] Figure 3 A schematic diagram of a directed acyclic graph of data dependency relationships between entities according to an embodiment of the present application.

[0017] Figure 4 A schematic diagram of a directed acyclic graph of deployment dependencies between entities according to an embodiment of the present application.

[0018] Figure 5 This is a flow chart of a data migration method according to an embodiment of the present application.

[0019] Figure 6 This is a flow chart of a data migration method according to an embodiment of the present application.

[0020] Figure 7 This is a structural block diagram of a data migration device according to an embodiment of the present application.

[0021] Figure 8 A block diagram of an electronic device used to implement an embodiment of the present application. DETAILED DESCRIPTION

[0022] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0023] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.

[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0025] Figure 1 This is a schematic diagram of an application scenario of the data migration method provided in this application. Figure 1 As shown, embodiments of the present application can be implemented as a business data warehouse migration and deployment device, where the current data warehouse environment is the source database and the data warehouse environment to be deployed is the target database. Objects to be migrated include code, configuration information, and scripts for processing data within a code development platform or database. A business data warehouse includes a data storage and management system used to support enterprise decision-making and business analysis. Business data warehouses typically utilize a variety of data storage and processing technologies, such as relational databases.

[0026] The business data warehouse migration and deployment device includes a business data warehouse deployment package generation module and a business data warehouse deployment package deployment module. The business data warehouse deployment package generation module is used to collect metadata, generate a directed acyclic graph (DAG) of metadata deployment relationships, extract metadata variables, and generate deployment packages. The business data warehouse deployment package deployment module is used for deployment preparation, configuration of deployment strategies, and topological sorting deployment.

[0027] Among them, metadata collection specifically includes: using open application programming interfaces to collect metadata of objects to be migrated from multiple cloud products in the current data warehouse environment. Metadata includes data that describes the data in the data warehouse and its related attributes, structure and relationships.

[0028] The metadata deployment relationship DAG generation specifically includes: using a type definition file to define the entity type of the object to be migrated based on the metadata. Entity types include any of the following: synchronization nodes, batch computing nodes, stream computing nodes, resources (scripts or data files that implement function functions), functions, data sources, data tables, etc. Entity dependency information is extracted from the metadata; this dependency information includes at least one of the entity's source data table information, the entity's destination data table information, the entity's dependent functions, or the entity's dependent data connectors. Data connectors include plugins that read and write data.

[0029] In a specific example, the entity and dependency information are as follows: Figure 2 As shown, data source 2 depends on Table 2 and Table 3, which are the destination data tables of data source 2; data source 1 depends on Table 1, which is the destination table of data source 1; synchronization node 3 depends on Table 1 and Table 2, which is the source data table of synchronization node 3 and Table 2 is the destination data table of synchronization node 3.

[0030] Based on dependency information, determine the entity's associated entities (such as Figure 2 As shown, the associated entities of data source 2 are Table 2 and Table 3); the entities and associated entities are taken as vertices, and the dependencies between entities are taken as edges between vertices to obtain a dependency graph; the circular dependency vertices are deleted from the dependency graph to obtain a directed acyclic graph; the circular dependency vertices are vertices that form a cyclic closed loop. Optionally, a graph traversal algorithm (such as a depth-first search algorithm) is used to detect whether there is a cycle in the dependency graph. By recording the traversed vertices and the vertices on the current traversal path, if a vertex is found to have appeared on the current traversal path during the traversal process, then the vertex is a circular dependency vertex, which is marked and deleted in the metadata.

[0031] In a specific example, the directed acyclic graph of data dependencies between entities is as follows: Figure 3 shown. Figure 3The data dependency relationships between entities in development platform 1 (synchronization node 3, resource 1, function 1, batch computing node 1) and entities in development platform 2 (stream computing node 2) are shown in FIG. The priority of the deployment order of the entities is determined according to the entity type. For example, the deployment order of the data table and the dependency information of the data table has a higher priority than that of the synchronization node, batch computing node, stream computing node, etc. According to the priority, the data dependency relationships between the entities are converted and processed to obtain the deployment dependency relationships between the entities. In a specific example, the directed acyclic graph of the data dependency relationships between the entities is as follows: Figure 4 shown. Figure 4 , which shows the deployment dependency relationship between entities in development platform 1 (synchronization node 3, resource 1, function 1, batch computing node 1) and entities in development platform 2 (stream computing node 2).

[0032] Metadata variable extraction specifically includes: obtaining the environment variables of the object to be migrated in the source data warehouse and saving them as a configuration file. The environment variables include at least one of the instance identifier (for example, database name, etc.), project identifier, business identifier, or fields required for data source connection.

[0033] The data package generation specifically includes: generating deployment data packages based on entity types, deployment dependencies between entities and environment variables, such as Figure 1 The data warehouse deployment package shown in . Transfer the data warehouse deployment package across environments to the data warehouse environment to be deployed for deployment.

[0034] Deployment preparation specifically includes: configuring the deployment package, loading the entities and deployment dependencies in the deployment package into the target data warehouse, and replacing the environment variables in the source data warehouse with those in the target data warehouse. For example, replace the instance ID of the cloud product in the source data warehouse with a meaningful name, such as replacing the database name with a recommended data source or a search data source. Replace the project names of batch and stream computing nodes with offline or real-time data warehouse names.

[0035] Configure a deployment strategy, which can include vertex-specific dependency deployment, full deployment, partial deployment, or incremental deployment. Vertex-specific dependency deployment (also known as dependency subgraph deployment, where a subgraph is a part of a graph consisting of a set of vertices and corresponding edges. In graph computing, a subgraph can represent a specific part or function) involves searching for the dependencies of a specified vertex, finding all dependent entities, and forming a dependency subgraph. Deployment is then performed based on the dependency subgraph. For a table-like vertex, the next-level vertex is first found, and then all dependent entities of the vertex are found to form a dependency subgraph. This deployment strategy only requires specifying a small amount of output data for a business. Based on the deployment dependencies, the business data warehouse can be fully deployed. This strategy is suitable for migration scenarios involving sub-businesses. Full deployment deploys all data warehouse objects and is suitable for overall migration. Partial deployment deploys specific deployment objects. Incremental deployment involves deploying based on metadata modification time. Configuring a specific modification time enables incremental metadata deployment and is suitable for incremental migration scenarios.

[0036] The topological sorting deployment specifically includes: outputting a list of entities to be deployed and a topological sorting based on the DAG of deployment dependencies and the deployment strategy; and deploying the deployment dependencies between entities to the target data warehouse based on the list of entities to be deployed and the topological sorting.

[0037] In this embodiment, open APIs are used to collect metadata from deployed objects across various cloud products that the business data warehouse relies on. This allows for migration across multiple cloud products and databases, supporting full or partial metadata migration of migrated objects. The entities of the migrated objects and the dependencies between them can be automatically extracted into a data deployment package and imported into the target data warehouse for rapid migration. This enables graph-based automatic extraction and automatic parameter extraction and configuration, ensuring the accuracy and completeness of the migration. This improves the efficiency and quality of data warehouse migration while ensuring data accuracy and integrity.

[0038] The embodiment of the present application provides a data migration method, which can be applied to computing devices, which may include servers, terminal devices, etc. Figure 5 The flowchart of the data migration method according to one embodiment of the present application is shown, including:

[0039] Step S501: Obtain metadata of the object to be migrated from the source data warehouse.

[0040] The objects to be migrated are the code, configuration information, and scripts used to process data in the code development platform or database. Metadata includes data that describes the data in the data warehouse and its related attributes, structure, and relationships.

[0041] Step S502: determining the entity type of the object to be migrated and the deployment dependency relationship between the entities based on the metadata.

[0042] Use the type definition file to define the entity type of the object to be migrated based on metadata. Entity types include synchronization nodes, batch computing nodes, streaming computing nodes, resources (scripts or data files that implement function functions), functions, data sources, and data tables.

[0043] Step S503: Obtain the environment variables of the object to be migrated in the source data warehouse.

[0044] Exemplarily, the environment variable includes at least one of an instance identifier, a project identifier, a business identifier, or a field required for a data source connection.

[0045] Step S504 : Migrating the objects to be migrated to the target data warehouse based on the entity types, deployment dependencies between entities, and environment variables.

[0046] The data migration method provided in an embodiment of the present application first obtains metadata of the object to be migrated from the source data warehouse; then, based on the metadata, determines the entity type of the object to be migrated and the deployment dependency relationship between entities; obtains the environment variables of the object to be migrated in the source data warehouse; and finally, based on the entity type, the deployment dependency relationship between entities, and the environment variables, migrates the object to be migrated to the target data warehouse. In an embodiment of the present application, the entity type of the object to be migrated, the deployment dependency relationship between entities, and the environment variables are obtained from the source data warehouse, and the object to be migrated is migrated to the target data warehouse based on the entity type, the deployment dependency relationship between entities, and the environment variables. This can achieve rapid data migration, improve the efficiency and quality of data migration, and ensure the accuracy and integrity of the data.

[0047] The following describes the specific implementation process of each of the above steps through various specific implementation methods:

[0048] In one implementation, in step S502, determining the entity type of the object to be migrated and the deployment dependency between the entities based on the metadata includes: determining the entity type of the object to be migrated and the data dependency between the entities based on the metadata; and determining the deployment dependency between the entities based on the entity type and the data dependency between the entities.

[0049] In practical applications, metadata is analyzed and type definition files are used to determine the entity types of the objects to be migrated. The upstream and downstream relationships between entities are analyzed to obtain data dependencies between entities. Based on the entity types, these data dependencies are converted into deployment dependencies.

[0050] In this embodiment, the deployment dependency relationship between entities is determined based on the entity type and the data dependency relationship between the entities, which can avoid the deployment dependency default problem.

[0051] In one implementation, data dependency relationships between entities are determined based on metadata, including: extracting dependency information of the entities from the metadata; and constructing a directed acyclic graph of dependency relationships between entities based on the entities and the dependency information, wherein the data dependency relationships include the directed acyclic graph.

[0052] The dependency information includes at least one of the source data table information of the entity, the destination data table information of the entity, the function that the entity depends on, or the data connector that the entity depends on.

[0053] The data connector includes plug-ins for reading and writing data.

[0054] In one implementation, a directed acyclic graph of dependency relationships between entities is constructed based on entity and dependency information, including: determining associated entities of an entity based on the dependency information; taking the entity and the associated entity as vertices, and the dependency relationships between the entities as edges between vertices to obtain a dependency graph; and deleting circular dependency vertices in the dependency graph to obtain a directed acyclic graph.

[0055] A circular dependency refers to a closed-loop dependency relationship between one or more vertices in a directed graph. A circular dependency vertex is a vertex that forms a closed loop.

[0056] Optionally, a graph traversal algorithm (such as a depth-first search algorithm) can be used to detect cycles in the dependency graph. This is done by recording the vertices that have been traversed and the vertices on the current traversal path. If a vertex is found to have already appeared on the current traversal path during the traversal process, the vertex is considered a circular dependency vertex and is marked and deleted in the metadata.

[0057] In one implementation, the deployment dependency between entities is determined based on the entity type and the data dependency between the entities, including: determining the priority of the deployment order of the entities based on the entity type; and converting the data dependency between the entities based on the priority to obtain the deployment dependency between the entities.

[0058] The deployment order priority can be set based on specific needs. The higher the priority, the higher the order in which the entity is deployed. For example, the deployment order priority of a data table and its dependencies is higher than that of a synchronization node, a batch computing node, a stream computing node, etc.

[0059] According to the priority, the directed acyclic graph of the data dependency relationships between the entities is transformed to obtain a directed acyclic graph of the deployment dependency relationships between the entities.

[0060] In one implementation, objects to be migrated are migrated to a target data warehouse based on entity types, deployment dependencies between entities, and environment variables, including: generating a deployment data package based on entity types, deployment dependencies between entities, and environment variables; and deploying the deployment data package to the target data warehouse.

[0061] The multiple entities to be migrated are organized based on entity types and deployment dependencies between entities, and together with environment variables, a deployment data package is generated and deployed to the target data warehouse.

[0062] In this embodiment, the entities of the objects to be migrated, the dependencies between the entities, and the environment variables can be automatically extracted into a data deployment package and imported into the target data warehouse to achieve rapid migration.

[0063] In one implementation, deploying a deployment data package to a target data warehouse includes: obtaining a deployment strategy; the deployment strategy includes specifying any one of vertex dependency deployment, full deployment, partial deployment, or incremental deployment; and deploying entities in the deployment data package and deployment dependency relationships between entities to the target data warehouse according to the deployment strategy.

[0064] Among them, the specified vertex dependency deployment (also known as dependency subgraph deployment, a subgraph refers to a part of a graph, which contains a set of vertices and corresponding edges. In graph computing, a subgraph can be used to represent a specific part or function): through the specified vertex, search for the dependency relationship of the vertex, find all dependent entities, form a dependency subgraph, and deploy according to the dependency subgraph. Among them, if the specified vertex is a data table, you can first find the next-level vertex, and then find all dependent entities of the vertex to form a dependency subgraph. This deployment strategy only needs to specify a small amount of output data information of a certain business. According to the deployment dependency relationship, the complete deployment of the business data warehouse can be carried out. It is suitable for deployment scenarios based on sub-business migration.

[0065] Full deployment involves deploying all data warehouse objects and is suitable for overall migration. Partial deployment involves deploying specific deployment objects. Incremental deployment involves deploying based on metadata modification time. By configuring a specific modification time, you can achieve incremental metadata deployment and is suitable for incremental migration scenarios.

[0066] In one implementation, deploying a deployment data package to a target data warehouse includes: loading the deployment data package into the target data warehouse and determining a topological sorting of entities based on deployment dependencies; and deploying entities in the deployment data package and deployment dependencies between entities to the target data warehouse based on the topological sorting.

[0067] Among them, topological sorting is a method for sorting vertices in a directed acyclic graph. Topological sorting can be used to determine the execution order of vertices in a directed acyclic graph.

[0068] Optionally, based on the DAG of the deployment dependency and the deployment strategy, a list of entities to be deployed and a topological sort are output; based on the list of entities to be deployed and the topological sort, the deployment dependency relationships between the entities are deployed to the target data warehouse.

[0069] In this embodiment, topological sorting is used to determine the deployment order of entities, which can ensure stable execution of the deployment process.

[0070] In one implementation, deploying the deployment package to the target data warehouse includes replacing environment variables in the source data warehouse in the deployment package with environment variables in the target data warehouse.

[0071] For example, replace the instance ID of the data source in the cloud product in the source data warehouse with a business-meaningful name, such as replacing the database name with "recommended data source" or "search data source." Replace the project name of the batch computing node and stream computing node with "offline data warehouse" or "real-time data warehouse."

[0072] The embodiment of the present application provides a data migration method, which can be applied to computing devices, which may include servers, terminal devices, etc. Figure 6 The flowchart of the data migration method according to one embodiment of the present application is shown, including:

[0073] Step S601: Obtain metadata of the object to be migrated from the source data warehouse.

[0074] The objects to be migrated are the code, configuration information, and scripts used to process data in the code development platform or database. Metadata includes data that describes the data in the data warehouse and its related attributes, structure, and relationships.

[0075] Step S602: determining the entity type of the object to be migrated and the data dependency relationship between the entities based on the metadata.

[0076] Use the type definition file to define the entity types of the objects to be migrated based on metadata. Entity types include synchronization nodes, batch computing nodes, stream computing nodes, resources (scripts or data files that implement function functions), functions, data sources, and data tables. Analyze the upstream and downstream relationships between entities to determine their data dependencies.

[0077] Step S603: Determine the priority of the deployment sequence of the entities according to the entity type.

[0078] The deployment order priority can be set based on specific needs. The higher the priority, the higher the order in which the entity is deployed. For example, the deployment order priority of a data table and its dependencies is higher than that of a synchronization node, a batch computing node, a stream computing node, etc.

[0079] Step S604: converting the data dependency relationships between entities according to the priorities to obtain deployment dependency relationships between the entities.

[0080] According to the priority, the directed acyclic graph of the data dependency relationships between the entities is transformed to obtain a directed acyclic graph of the deployment dependency relationships between the entities.

[0081] Step S605: Obtain the environment variables of the object to be migrated in the source data warehouse.

[0082] Exemplarily, the environment variable includes at least one of an instance identifier, a project identifier, a business identifier, or a field required for a data source connection.

[0083] Step S606: Generate a deployment data package based on the entity type, deployment dependencies between entities, and environment variables.

[0084] The multiple entities to be migrated are organized based on entity types and deployment dependencies between entities, and together with environment variables, a deployment data package is generated.

[0085] Step S607: Obtain the deployment strategy.

[0086] The deployment strategy includes specifying any one of vertex dependency deployment, full deployment, partial deployment, or incremental deployment.

[0087] Among them, the specified vertex dependency deployment (also known as dependency subgraph deployment, a subgraph refers to a part of a graph, which contains a set of vertices and corresponding edges. In graph computing, a subgraph can be used to represent a specific part or function): through the specified vertex, search for the dependency relationship of the vertex, find all dependent entities, form a dependency subgraph, and deploy according to the dependency subgraph. Among them, if the specified vertex is a data table, you can first find the next-level vertex, and then find all dependent entities of the vertex to form a dependency subgraph. This deployment strategy only needs to specify a small amount of output data information of a certain business. According to the deployment dependency relationship, the complete deployment of the business data warehouse can be carried out. It is suitable for deployment scenarios based on sub-business migration.

[0088] Full deployment involves deploying all data warehouse objects and is suitable for overall migration. Partial deployment involves deploying specific deployment objects. Incremental deployment involves deploying based on metadata modification time. By configuring a specific modification time, you can achieve incremental metadata deployment and is suitable for incremental migration scenarios.

[0089] Step S608: Determine the topological order of the entities according to the deployment dependency.

[0090] Among them, topological sorting is a method for sorting vertices in a directed acyclic graph. Topological sorting can be used to determine the execution order of vertices in a directed acyclic graph.

[0091] Step S609: deploy the deployment data package to the target data warehouse according to the deployment strategy and topological sorting.

[0092] According to the directed acyclic graph of deployment dependencies and the deployment strategy, a list of entities to be deployed and a topological sort are output; according to the list of entities to be deployed and the topological sort, the deployment dependencies between entities are deployed to the target data warehouse.

[0093] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a data migration device. Figure 7 FIG2 is a block diagram of a data migration device according to an embodiment of the present application, which includes:

[0094] The first acquisition module 701 is used to obtain metadata of the object to be migrated from the source data warehouse;

[0095] Determining module 702, for determining the entity type of the object to be migrated and the deployment dependency relationship between the entities based on the metadata;

[0096] The second acquisition module 703 is used to obtain the environment variables of the object to be migrated in the source data warehouse;

[0097] The migration module 704 is used to migrate the objects to be migrated to the target data warehouse based on entity types, deployment dependencies between entities, and environment variables.

[0098] The data migration device provided in an embodiment of the present application first obtains metadata of the object to be migrated from the source data warehouse; then, based on the metadata, determines the entity type of the object to be migrated and the deployment dependency relationship between entities; obtains the environment variables of the object to be migrated in the source data warehouse; and finally, based on the entity type, the deployment dependency relationship between entities, and the environment variables, migrates the object to be migrated to the target data warehouse. In an embodiment of the present application, the entity type of the object to be migrated, the deployment dependency relationship between entities, and the environment variables are obtained from the source data warehouse, and the object to be migrated is migrated to the target data warehouse based on the entity type, the deployment dependency relationship between entities, and the environment variables, which can achieve rapid data migration, improve the efficiency and quality of data migration, and ensure the accuracy and integrity of the data.

[0099] In one implementation, the determination module 702 is configured to: determine the entity type of the object to be migrated and the data dependency relationship between the entities based on the metadata; and determine the deployment dependency relationship between the entities based on the entity type and the data dependency relationship between the entities.

[0100] In one implementation, the determination module 702 is used to: extract the dependency information of the entity from the metadata when determining the data dependency relationship between entities based on the metadata; construct a directed acyclic graph of the dependency relationship between the entities based on the entities and the dependency information, and the data dependency relationship includes a directed acyclic graph; wherein the dependency information includes: at least one of the source data table information of the entity, the destination data table information of the entity, the function on which the entity depends, or the data connector on which the entity depends.

[0101] In one implementation, the determination module 702 is used to: when constructing a directed acyclic graph of dependency relationships between entities based on entities and dependency information, determine the associated entities of the entity based on the dependency information; use the entity and the associated entity as vertices, and the dependency relationships between the entities as edges between the vertices to obtain a dependency graph; delete the circular dependency vertices in the dependency graph to obtain a directed acyclic graph; the circular dependency vertices are vertices that form a cyclic closed loop.

[0102] In one implementation, the determination module 702 is used to: when determining the deployment dependency between entities based on the entity type and the data dependency between the entities, determine the priority of the deployment order of the entities based on the entity type; and convert the data dependency between the entities based on the priority to obtain the deployment dependency between the entities.

[0103] In one implementation, the migration module 704 is configured to: generate a deployment data package based on entity types, deployment dependencies between entities, and environment variables; and deploy the deployment data package to a target data warehouse.

[0104] In one implementation, the migration module 704 is used to: obtain a deployment strategy when deploying a deployment data package to a target data warehouse; the deployment strategy includes specifying any one of vertex dependency deployment, full deployment, partial deployment, or incremental deployment; and deploy entities in the deployment data package and deployment dependencies between entities to the target data warehouse according to the deployment strategy.

[0105] In one implementation, the migration module 704 is used to: when deploying a deployment data package to a target data warehouse, load the deployment data package into the target data warehouse and determine the topological sorting of the entities based on the deployment dependencies; and deploy the entities in the deployment data package and the deployment dependencies between the entities to the target data warehouse based on the topological sorting.

[0106] In one implementation, the migration module 704 is used to: when deploying a deployment package to a target data warehouse, replace the environment variables in the source data warehouse in the deployment package with the environment variables in the target data warehouse, where the environment variables include at least one of an instance identifier, a project identifier, a business identifier, or a field required for a data source connection.

[0107] The functions of each module in the embodiment of the present application can be referred to the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0108] Figure 8 FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present application. Figure 8 As shown, the electronic device includes: a memory 810 and a processor 820. The memory 810 stores a computer program that can be run on the processor 820. When the processor 820 executes the computer program, the method in the above embodiment is implemented. The number of the memory 810 and the processor 820 can be one or more.

[0109] The electronic device also includes:

[0110] The communication interface 830 is used to communicate with external devices and perform data exchange transmission.

[0111] If the memory 810, the processor 820, and the communication interface 830 are implemented independently, the memory 810, the processor 820, and the communication interface 830 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0112] Optionally, in a specific implementation, if the memory 810, the processor 820 and the communication interface 830 are integrated on a chip, the memory 810, the processor 820 and the communication interface 830 can communicate with each other through an internal interface.

[0113] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present application when the program is executed by a processor.

[0114] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.

[0115] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0116] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0117] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).

[0118] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0119] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0120] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0121] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.

[0122] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.

[0123] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0124] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0125] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A data migration method, characterized in that: The method comprises: Obtain metadata of the objects to be migrated from the source data warehouse; Determining, based on the metadata, entity types of the objects to be migrated and deployment dependencies between entities; Obtaining environment variables of the object to be migrated in the source data warehouse; The objects to be migrated are migrated to a target data warehouse based on the entity types, the deployment dependencies between the entities, and the environment variables.

2. The method according to claim 1, characterized in that Determining the entity type of the object to be migrated and the deployment dependency relationship between entities based on the metadata includes: Determining, based on the metadata, entity types of the objects to be migrated and data dependencies between entities; Determine the deployment dependency relationship between the entities according to the entity types and the data dependency relationship between the entities.

3. The method according to claim 2, characterized in that Determine data dependencies between entities based on the metadata, including: extracting dependency information of the entity from the metadata; Based on the entities and the dependency information, constructing a directed acyclic graph of dependency relationships between entities, wherein the data dependency relationship includes the directed acyclic graph; The dependency information includes at least one of the source data table information of the entity, the destination data table information of the entity, the function on which the entity depends, or the data connector on which the entity depends.

4. The method according to claim 3, characterized in that The step of constructing a directed acyclic graph of dependency relationships between entities based on the entities and the dependency information includes: Determining an associated entity of the entity based on the dependency information; Taking the entity and the associated entity as vertices and the dependency relationships between the entities as edges between the vertices, a dependency graph is obtained; The circular dependency vertex is deleted from the dependency graph to obtain the directed acyclic graph; the circular dependency vertex is a vertex that forms a circular closed loop.

5. The method according to claim 2, characterized in that The determining, based on the entity types and the data dependencies between the entities, the deployment dependencies between the entities includes: Determining the priority of the deployment order of the entities according to the entity type; According to the priorities, the data dependencies between the entities are converted to obtain deployment dependencies between the entities.

6. The method according to claim 1, characterized in that The migrating the object to be migrated to the target data warehouse based on the entity type, the deployment dependency relationship between the entities, and the environment variable includes: generating a deployment data package based on the entity type, the deployment dependency relationship between the entities, and the environment variables; Deploy the deployment data package to the target data warehouse.

7. The method according to claim 6, characterized in that The deploying the deployment data package to the target data warehouse includes: Obtaining a deployment strategy; the deployment strategy includes any one of specifying vertex-dependent deployment, full deployment, partial deployment, or incremental deployment; According to the deployment strategy, the entities in the deployment data package and the deployment dependency relationships between the entities are deployed to the target data warehouse.

8. The method according to claim 6, characterized in that The deploying the deployment data package to the target data warehouse includes: Loading the deployment data package into the target data warehouse, and determining a topological sorting of the entities based on the deployment dependencies; According to the topological sorting, the entities in the deployment data package and the deployment dependency relationships between the entities are deployed to the target data warehouse.

9. The method according to any one of claims 1 to 8, characterized in that The deploying the deployment data package to the target data warehouse includes: The environment variables in the source data warehouse in the deployment data package are replaced with the environment variables in the target data warehouse, wherein the environment variables include at least one of an instance identifier, a project identifier, a business identifier, or a field required for data source connection.

10. A data migration device, characterized in that: The device comprises: A first acquisition module is used to acquire metadata of the object to be migrated from the source data warehouse; a determination module, configured to determine, based on the metadata, entity types of the objects to be migrated and deployment dependencies between entities; A second acquisition module is used to acquire the environment variables of the object to be migrated in the source data warehouse; A migration module is used to migrate the objects to be migrated to a target data warehouse based on the entity types, the deployment dependencies between the entities and the environment variables.

11. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory, and the processor implements the method according to any one of claims 1 to 9 when executing the computer program.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Domestic Gaussian database substitution method for electricity utilization information acquisition system

    CN121326888A