Data cutover method and device
By obtaining the separating rules of the target computing medium, and automatically determining the execution order of the data conversion task, the problem of lack of standardization and universality of the data separating scheme in the prior art is solved, and the efficiency and stability of data separating are achieved.
Patent Information
- Application Number
- CN202110326458.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-26
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-03-26
AI Technical Summary
Existing data separating schemes lack standardization and versatility, and execution efficiency and stability rely on manual experience and cannot adapt to different computing media.
Provide a data separating method and device, by obtaining the separating rules of the target computing medium, automatically determine the execution order of the data conversion task, and perform data conversion in the target computing medium, and finally transmit the target data to the target end.
The standardization and generalization of the data separating process are realized, execution efficiency and stability are improved, and manual intervention is reduced.
Smart Images

Figure CN115129439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data engineering technology, and in particular to a data cutover method and device. Background Art
[0002] With the rise in popularity of cloud computing, more and more customers are building their systems in the cloud. Furthermore, with this increasing adoption, more and more users are increasingly looking to migrate their core off-premises systems to the cloud. Unlike migrating new systems to the cloud, core system migration often requires compatibility with existing data. To ensure that the new system can accurately utilize data from the old system, the old system's data must be converted according to pre-defined rules and then migrated to the new system during the switchover. This data conversion and migration is called data cutover.
[0003] Currently, the industry lacks a standardized data cutover solution for the specific steps involved. Instead, solutions are often customized based on specific projects and the computing media used for the cutover, making them extremely incompatible. Furthermore, the execution order of cutover tasks in existing cutover solutions is often manually determined based on experience, making the efficiency of the cutover dependent on manual experience and significantly impacting its efficiency and stability. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a data cutover method and apparatus, the main purpose of which is to provide a standardized data cutover solution that is universally applicable to various computing media.
[0005] To achieve the above objectives, the present invention mainly provides the following technical solutions:
[0006] In a first aspect, the present invention provides a data cutover method, applied to a target computing medium, wherein the target computing medium is used to convert source data into target data, the method comprising:
[0007] Use data transfer services to obtain source data from the source;
[0008] Obtaining a cutover rule corresponding to the source data and applicable to the target computing medium;
[0009] Determining a data conversion task having an execution order according to the first association relationship of the source data and the cutover rule, wherein the data conversion task is a conversion operation corresponding to converting the source data having the first association relationship into target data having the second association relationship according to the cutover rule;
[0010] Execute the data conversion task according to the execution order to obtain target data having a second association relationship;
[0011] The target data is transmitted to the target end.
[0012] In a second aspect, the present invention provides a data cutover device, applied to a target computing medium, wherein the target computing medium is used to convert source data into target data, the device comprising:
[0013] A first transmission unit, configured to obtain source data from a source end using a data transmission service;
[0014] an acquiring unit, configured to acquire a cutover rule corresponding to the source data obtained by the first transmission unit and applicable to the target computing medium;
[0015] a determining unit, configured to determine a data conversion task having an execution order based on the first association relationship of the source data and the cutover rule obtained by the acquiring unit, wherein the data conversion task is a conversion operation corresponding to converting the source data having the first association relationship into target data having the second association relationship according to the cutover rule;
[0016] a conversion unit, configured to execute the data conversion task determined by the determination unit according to the execution order, to obtain target data having a second association relationship;
[0017] The second transmission unit is configured to transmit the target data obtained by the conversion unit to a target end.
[0018] In a third aspect, the present invention provides a data cutover system, comprising: a source end, a data cutover end, and a target end; wherein the data cutover end is applied to a target computing medium, and the target computing medium is used to convert source data into target data;
[0019] The data cutover end is configured to call a data transmission service according to a cutover request to obtain source data from a source end, convert the source data into target data based on the data cutover method described in the first aspect, and transmit the target data to the target end;
[0020] The source end is used to transmit the source data corresponding to the cutover request to the target computing medium corresponding to the data cutover end by using a data transmission service.
[0021] In another aspect, the present invention provides a processor, which is configured to run a program, wherein the program executes the above-mentioned data cutover method when running.
[0022] On the other hand, the present invention provides a computer-readable storage medium for storing a computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned data cutover method.
[0023] Through the above technical solution, the present invention provides a data cutover method and apparatus that are applicable to a variety of different computing media and propose a standardized solution for the data cutover process. This solution transfers source data from a source end to a target computing medium, obtains corresponding cutover rules from the target computing medium, and uses these cutover rules in combination with the first association relationship of the source data to resolve the data conversion tasks to be executed. The solution also automatically determines the execution order between the tasks and converts the source data according to the execution order to obtain target data with a second association relationship. Finally, the target data is transferred from the target computing medium to the target end, completing the data cutover. Compared to existing data cutover methods, the data cutover process of the present invention only needs to consider providing cutover rules applicable to the target computing medium. The generation and execution order of data conversion tasks are independent of the target computing medium. As a result, this part can be reused across different computing media, achieving a universal data cutover process. Furthermore, because the embodiments of the present invention can automatically schedule data conversion tasks without manual planning, the data cutover solution can also effectively improve the efficiency and stability of data cutover execution.
[0024] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0026] Figure 1 A schematic diagram of a data cutover process framework according to an embodiment of the present invention is shown;
[0027] Figure 2 A flow chart of a data cutover method proposed in an embodiment of the present invention is shown;
[0028] Figure 3 A flow chart of another data cutover method proposed in an embodiment of the present invention is shown;
[0029] Figure 4 A schematic diagram showing the steps of a data cutover process according to an embodiment of the present invention is shown;
[0030] Figure 5 The corresponding embodiment of the present invention is shown Figure 4 Schematic diagram of the constructed directed acyclic graph;
[0031] Figure 6 A block diagram showing the composition of a data cutover device proposed in an embodiment of the present invention is shown;
[0032] Figure 7 The figure shows a block diagram of another data cutover device proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0034] The data cutover method proposed in the present invention is based on the popularity of cloud computing. It is a universal and standardized data migration solution proposed to meet the needs of users to migrate their core systems from the cloud to the cloud. In this solution, the data migration and conversion is mainly performed by the data cutover end, and the data cutover end can adopt data cutover rules in corresponding formats according to the computing media used by the user. The computing media may include but is not limited to databases, application systems, big data platforms, etc.
[0035] The embodiment of the present invention takes the computing medium as a database as an example, that is, the data cutover end is applied to the database, and the data cutover process is as follows: Figure 1As shown, the source end in the figure is used to store the data of the core system under the cloud, while the target end is used to store the data of the system on the cloud. The cutover transit database in the figure serves as the computing medium used by the customer to perform data cutover operations to realize the conversion and migration of data from the source end to the target end. The specific process is as follows: the user triggers a data cutover request at the data cutover end. The data cutover request contains information on obtaining specified data from the specified source end. The data cutover end calls the data transmission service DTS (Data Transmission Service) according to the data cutover request to obtain the specified data from the specified source end and store the obtained data in the corresponding table (source database mirror) in the database; after that, the data cutover end will determine the first association relationship of the data based on the obtained data through the data model of the old system (core system under the cloud), and at the same time, obtain the cutover rules applicable to the database. The cutover rules are pre-set to convert the first association relationship of the data into the second association relationship applicable to the new system (core system on the cloud); in this way, the data cutover end can convert the obtained data according to the cutover rules, convert the source data with the first association relationship into the target data with the second association relationship, and save the target data in the database In the corresponding position (target library mirror), it should be noted that in the data conversion process, the association relationship between the data is not one-to-one. Therefore, when performing cutover operations on a large amount of data, in order to improve the cutover efficiency, it is necessary to optimize the sequence of different data cutover operations. Compared with the existing manually formulated execution sequence, the present invention automatically optimizes the execution sequence of different data based on the analysis of the cutover rules and source data, generates data conversion tasks with an execution sequence, and executes them in sequence to obtain target data with a second association relationship; finally, the data cutover end transmits the converted target data to the target end through a data transmission tool (the data synchronization tool DataX shown in the figure). Based on the description of the above process, the data cutover process can also be reused in other computing media, and the difference in its reuse process is only that it is necessary to obtain the cutover rules applicable to the corresponding computing medium.
[0036] Based on the above description of the overall process of the present invention, the specific implementation of each step in a data cutover method provided by an embodiment of the present invention will be described in detail below through specific embodiments. The specific steps of this method are as follows: Figure 2 As shown, the method includes:
[0037] Step 101: Obtain source data from a source end using a data transmission service.
[0038] The data cutover method implemented in this embodiment depends on the target computing medium used to convert source data into target data, i.e., the computing medium used for data conversion during data cutover. Common computing media include databases, application systems, and big data platforms. Due to the significant performance differences between different computing media, the order of executing conversion tasks must be determined based on the performance of the computing media when determining data conversion. This requires customized data cutover solutions for different computing media, making existing data cutover solutions non-universal.
[0039] The present invention provides a data cutover solution applicable to various computing media. Therefore, the target computing medium (i.e., the target computing medium) must be identified. Based on this computing medium, a corresponding data transfer service is invoked to import the source data into the target computing medium. This data transfer service is not limited to a dedicated data transfer service for the source data type, and can also be applied to multiple data types.
[0040] Step 102: Obtain a cutover rule corresponding to the source data and applicable to the target computing medium.
[0041] The migration rule is a manually formulated data conversion rule based on the first association between the source data and the second association between the target data. In this embodiment, the first and second associations can be the same or different. For example, if the amount data is associated with the product name, and the amount in the source data is in yuan, while the amount in the migrated target data is in cents, the corresponding migration rule could be to multiply the amount in the source data by 100, resulting in the amount data as the target data. The second association of the target data remains between amount and product name. Clearly, the migration rule in this example only processes the data content and does not change the associations. For another example, if the source data is Table A (product code-product name) and Table B (product code-product amount), and the migration rule is to obtain Table X, which maps product names and amounts based on product codes, then the target data would be Table X (product name-product amount). Clearly, in this example, the migration rule maps the first associations in the source data (product code-product name, product code-product amount) to the second association (product name-product amount) by associating the product codes.
[0042] Furthermore, because different computing media require different cutover rule formats, obtaining cutover rules requires obtaining cutover rules that are recognizable by the target computing media. For example, for a database, the cutover rule format might be a script written in the structured query language SQL, while for a big data platform, such as the open-source cluster computing environment Spark, the cutover rule format might be a script written in the Scala programming language. Therefore, these cutover rules require technical personnel to customize and develop them based on actual needs.
[0043] In actual project execution, data cutover rules are often not fixed but rather subject to change and update based on real-time needs. Therefore, in this embodiment, the step of obtaining cutover rules is a real-time operation. Generally, this can be done by passively obtaining updated cutover rules based on uploads from rule developers, or by periodically polling the address where cutover rules are stored to determine if they have been updated. If a new cutover rule is found, subsequent data cutover operations are performed based on that new rule. Thus, through this cutover rule acquisition step, this solution provides the ability to continuously update cutover rules.
[0044] In addition, with respect to obtaining cutover rules applicable to the target computing medium, this embodiment does not specifically limit the cutover rules to be written independently using code recognizable by the target computing medium, or to be edited using a cutover rule template provided by the target computing medium.
[0045] Step 103: Determine data conversion tasks with an execution order according to the first association relationship of the source data and the cutover rule.
[0046] This step is the core step in the data cutover process. It is to atomically split the source data according to the first association relationship to obtain atomic data that cannot be further split. Then, the association relationship between these atomic data is reconstructed according to the cutover rules, and the atomic data is reorganized according to the second association relationship to obtain the target data. However, during the process of splitting and reorganizing the source data, because the same atomic data often has a second association relationship with multiple other atomic data according to the cutover rules, it is necessary to perform multiple conversion operations when performing the cutover conversion operation on the atomic data. Due to the limited parallel computing capabilities of the computing medium, it is necessary to set a corresponding execution order for these conversion operations.
[0047] In this step, a data conversion task refers to the conversion operation corresponding to converting source data with a first association into target data with a second association according to the cutover rules. When a conversion operation has multiple steps, multiple data conversion tasks may be associated with it. That is, the target data is generated based on multiple data conversion tasks with associated relationships. The process of converting one source data into one target data typically involves multiple data conversion tasks, and the execution order of these data conversion tasks directly affects the efficiency of the data conversion. Therefore, this embodiment, based on analysis of the cutover rules, can plan the execution order of data conversion tasks. This execution order includes both parallel and sequential execution. In other words, the ideal state for executing a data conversion task is that when a data conversion task is executed, the other data conversion tasks it depends on have already completed, without having to wait for the execution results of the other data conversion tasks. Therefore, this step can utilize a task scheduling tool to plan the execution order of data conversion tasks according to the cutover rules and execute the data conversion tasks based on this execution order.
[0048] In this step, the execution of the data conversion task is independent of the properties of the target computing medium, while the cutover rule is adapted to the target computing medium. Therefore, the execution of this step is applicable to a variety of computing media.
[0049] Step 104: Execute the data conversion task according to the execution order to obtain target data with a second association relationship.
[0050] This step involves converting the source data in the target computing medium according to the execution order of the data conversion tasks to obtain target data having the second association relationship. Specifically, this step involves reasonably scheduling the execution order of the data conversion tasks determined in the previous step to improve the conversion efficiency of obtaining the target data.
[0051] Step 105: Transmit the target data to the target end.
[0052] This step is the final step in the data cutover process, publishing the converted target data from the target computing medium to the target end. In actual applications, this step can be done by transferring the target data through a pre-set data transfer tool.
[0053] Through the description of the above embodiments, a data cutover method provided by an embodiment of the present invention is a standardized data cutover process solution that can be applied to a variety of computing media. The solution includes importing source data into a target computing medium and converting the source data by obtaining a cutover rule applicable to the target computing medium. In the conversion process, the source data is parsed and divided into atomic data based on a first association relationship, and the atomic data is reorganized into target data with a second association relationship in combination with the cutover rule. For its segmentation and reorganization process, this embodiment uses the conversion operation required to convert the target data as the corresponding data conversion task, and improves the data cutover efficiency by optimizing the execution order of the data conversion task. Since the conversion process is independent of the computing medium, the process can be reused for different computing media, and the data conversion task can be executed in a determined order in different solution media. Finally, the target data obtained after conversion is transmitted from the target computing medium to the target end to complete the data cutover process. Compared to the existing approach of customizing and developing data cutover solutions for different computing media, the data cutover process proposed in the embodiment of the present invention is universal for multiple computing media. At the same time, this embodiment also has automatic scheduling of task execution order, eliminating the need for manual formulation. Therefore, this data cutover solution can also effectively improve the execution efficiency and stability of data cutover.
[0054] Further, for Figure 2 The data cutover method shown in the following embodiment will be described in detail using a relational database as an example. In this embodiment, the source end and the target end are relational databases, the source data and the target data are data tables in the relational database, and the computing medium is a database. In this application scenario, the specific process of the above data cutover solution is as follows: Figure 3 Shown, including:
[0055] Step 201: Obtain source data from a source end using a data transmission service.
[0056] Because the target computing medium in this embodiment is an intermediate database, and to accommodate different types of source databases, the data transmission service employed in this embodiment requires a tool for transferring different types of databases to the intermediate database, such as the DTS data transmission service provided by Alibaba Cloud. This transfer step isolates the impact on the source database, ensuring that the source database can function normally without being affected by the data cutover process.
[0057] Step 202: Obtain a cutover rule corresponding to the source data and applicable to the target computing medium.
[0058] The cutover rule script obtained in this step is applicable to the intermediate database. For relational databases, the data relationships within their data tables are generally represented by data models. Therefore, the data relationships in the old system correspond to the old data model, while the data relationships in the new system correspond to the new data model. Using these cutover rules, the old data model can be converted to the new data model. By using structured expressions such as mapping tables, SQL (Structured Query Language), and XML (Extensible Markup Language), the first relationships in the source data are converted to target data with the second relationships. It should be noted that in this embodiment, the data cutover process not only includes adjustments to data relationships but also changes to the data content.
[0059] Step 203: Split the source data according to the preset splitting rules to obtain atomic data for performing the conversion operation.
[0060] Specifically, for the splitting of data tables in relational databases, the preset splitting rules are pre-set splitting dimensions for specific data tables, for example, specified fields in the data tables. By splitting according to the preset dimensions, a data table can be split into multiple sub-data tables, and the first association relationship of the data in the data table is converted into the association relationship between the sub-data tables. After the splitting, the minimum step based on the cutover rule is obtained, that is, the atomic data. According to the cutover rule, the second association relationship of the data table in the new system can be understood as being obtained by performing operations such as association and mapping on the atomic data according to the cutover rule. In order to facilitate the execution of subsequent cutover rules in the intermediate database, the sub-data tables obtained after the splitting also need to be indexed. Of course, for other computing media, such as distributed computing platforms, there is no need to build indexes.
[0061] For the process of this step, please refer to Figure 4In the figure, A, B, C, and X represent four data tables in the source database. These tables are split according to preset dimensions, with table A correspondingly split into three sub-tables: a1, a2, and a3. Tables B and X are split according to the first association relationship to obtain sub-tables β and β*, where β* represents multiple different sub-tables. Table C does not need to be split. To facilitate association mapping of the sub-tables according to the cutover rules, indexes are constructed for the split sub-tables. Then, according to the obtained cutover rules, the sub-tables are associated and mapped, resulting in the target data sub-tables ab1, ab2, ab3, bc, and cc. After aggregation, auditing, and other processing, the new data tables AB, bc, and cc with the second association relationship are finally obtained on the target side. As shown in the figure, a1, a2, a3, β, β*, and c represent the atomic data after splitting. The resulting atomic data are the smallest granularity sub-tables for performing association mapping and other transformation operations according to the cutover rules.
[0062] Step 204: Construct a directed acyclic graph for describing the data conversion task based on the atomic data and the cutover rules.
[0063] The specific implementation process of this step includes: determining the dependency relationship between atomic data according to the cutover rules; constructing a directed acyclic graph with atomic data as nodes and dependency relationships as edges; and determining a group of atomic data with serial dependency relationships in the directed acyclic graph as a data conversion task.
[0064] based on Figure 4 The data cutover process shown in the figure can use the data tables as nodes and the dependency relationships of the generated data tables as edges to obtain Figure 5 In the directed acyclic graph shown, each node can be understood as being synthesized from other nodes according to the cutover rules. Thus, each node can be abstracted as a composite of multiple nodes, each of which can be considered a data conversion subtask. A data conversion task can be a collection of one or more subtasks. For example, subtask 1 corresponding to table β is the join operation between table X and table B, and subtask 2 corresponding to table c is the indexing of table C. Then, data conversion task 3 corresponding to table bc is the join between subtask 1 and subtask 2, and data conversion task 4 corresponding to table cc is subtask 2. Thus, the subtasks contained in data conversion task 5 corresponding to table AB and the subtasks contained in data conversion task 4 do not have a mutual dependency relationship. However, both data conversion task 5 and data conversion task 3 require the execution of subtask 1. Therefore, when the execution of data conversion tasks is subsequently scheduled, data conversion task 4 can be executed concurrently with data conversion tasks 3 and 5. However, for data conversion tasks 3 and 5, the execution order is to execute subtask 1 first, followed by the execution of the other subtasks.
[0065] Step 205: Determine the execution order of the data conversion tasks based on the data conversion task scheduling framework corresponding to the directed acyclic graph.
[0066] according to Figure 5 In the directed acyclic graph shown in FIG, when determining the order of each data conversion task and subtask, the existing task scheduling framework based on the directed acyclic graph can be used to plan the preferred scheduling execution order according to the relevant parameters of each subtask, including the concurrent and sequential execution schemes of the tasks. Commonly used task scheduling frameworks for directed acyclic graphs include Airflow, and this embodiment does not limit the scheduling framework tool used. That is, existing tools can be applied, and task scheduling tools designed by technicians based on directed acyclic graphs can also be applied.
[0067] In this embodiment, the data conversion task scheduling framework based on a directed acyclic graph not only schedules the execution order of each data conversion task, improving the efficiency of the data cutover process, but also performs functions such as concurrency control, time consumption analysis, blocking analysis, and power-off retry on the scheduled tasks, thereby better utilizing computing resources and ensuring the stability of task execution. Furthermore, these functions can be visualized as needed, helping execution personnel more intuitively view the real-time status of the data cutover process.
[0068] correspond Figure 4 As shown at the bottom of the data cutover process, it can be found that although specific operation steps are defined for the data cutover process in this embodiment, due to the addition of task scheduling of a directed acyclic graph, in the actual execution process, the conversion of different data tables is not necessarily performed strictly in the order of the steps, that is, all operations in one step are completed before entering the next step. Instead, when computing performance permits, the data conversion tasks are executed in the scheduled order.
[0069] Step 206: Execute the data conversion task according to the execution order to obtain target data with a second association relationship.
[0070] In this embodiment, the execution efficiency of data cutover is mainly reflected in the execution efficiency of the data conversion task, and the concurrency of the task is related to the execution efficiency. The larger the concurrency, the more tasks are executed simultaneously, and the higher the execution efficiency. Although the task scheduling tool of the directed acyclic graph mentioned above also has the ability of concurrency control, its concurrency control is for the scheduling tool and is not strongly related to the computing performance of the computing medium. To this end, in a preferred embodiment of the present invention, a dynamic control device can also be set for the target computing medium (especially for the intermediate database) for this step, which is used to dynamically monitor the computing performance of the computing medium when executing data cutover, adjust the concurrency of the data conversion task according to its monitoring results, and further provide real-time feedback on problems that arise during the execution process. In this way, by performing secondary scheduling on the executed data conversion tasks, the stability of data cutover execution can be further improved.
[0071] Furthermore, while performing secondary scheduling based on the computing performance of the computing medium, since the data transmission service in step 201 can transmit multiple types of data, when the source data transmitted from the source end to the target computing medium has a second data format, such as application data, and the source data in the second data format is not applicable to the above-mentioned conversion steps, for this reason, a preferred embodiment of the present invention can generate corresponding data conversion tasks for the source data in the second data format according to a preset strategy, and use a preset scheduler to schedule the execution order of the data conversion tasks, wherein the preset scheduler is used to cooperate with the data conversion task scheduling framework corresponding to the directed acyclic graph to determine the execution order of the data conversion tasks. By setting the scheduler, the types of data obtained from the source end can be expanded, thereby further improving the versatility of the data cutover solution.
[0072] Step 207: Audit the target data and source data using a data audit script.
[0073] This step involves performing a data audit after data conversion to improve the accuracy of the data cutover. This audit includes data aggregation, consistency checks, data relationship audits, and data logic audits. These audits can be automatically executed using a pre-set data audit script. To this end, before executing this step, it is necessary to obtain the data audit script. This data audit script can be obtained simultaneously with the cutover rules, and its execution can be coordinated and scheduled as an audit task, enabling real-time data audits once the target data required for the audit task is available.
[0074] Step 208: Transmit the target data having the second association relationship to the target end.
[0075] above Figure 3The illustrated embodiment illustrates a data cutover process using a database as the computing medium and a relational database as the source and target. The core steps of the cutover process are the generation and scheduling of data conversion tasks. As demonstrated in the preceding embodiment, the execution of this component is independent of the computing medium, making it reusable across multiple computing media. Consequently, the data cutover solution proposed by the present invention can be applied to various computing media. Furthermore, by utilizing a directed acyclic graph-based data conversion task scheduling framework to optimize the execution of data conversion tasks, compared to manually customizing task sequences, this not only reduces labor costs but also maximizes computing resource utilization, thereby improving data cutover efficiency. Furthermore, secondary task scheduling further ensures the stability of the data cutover within the computing medium, enhancing system security.
[0076] Furthermore, as a response to the above Figure 2 、 3 In order to realize the data cutover method shown in the figure, the embodiment of the present invention provides a data cutover device, the main purpose of which is to provide a standardized and feasible solution for data cutover operation that is applicable to various computing media. For the sake of ease of reading, this device embodiment will not repeat the details of the above method embodiment one by one, but it should be clear that the device in this embodiment can realize all the contents of the above method embodiment. Figure 6 As shown, it is mainly applied to the target computing medium, which is used to convert the source data into the target data, specifically including:
[0077] The first transmission unit 31 is configured to obtain source data from a source terminal using a data transmission service;
[0078] An acquiring unit 32 is configured to acquire a cutover rule corresponding to the source data obtained by the first transmission unit 31 and applicable to the target computing medium;
[0079] a determining unit 33 configured to determine a data conversion task having an execution order based on the first association relationship of the source data and the cutover rule obtained by the obtaining unit 32, wherein the data conversion task is a conversion operation corresponding to converting the source data having the first association relationship into target data having the second association relationship according to the cutover rule;
[0080] A conversion unit 34 is configured to execute the data conversion task determined by the determination unit 33 according to the execution order to obtain target data having a second association relationship;
[0081] The second transmission unit 35 is configured to transmit the target data obtained by the conversion unit 34 to a target end.
[0082] Further, such as Figure 7 As shown, the determining unit 33 includes:
[0083] A splitting module 331 is used to split the source data according to a preset splitting rule to obtain atomic data for performing a conversion operation;
[0084] a determination module 332 configured to construct a directed acyclic graph (DAG) for expressing a data conversion task based on the atomic data obtained by the splitting module 331 and the cutover rule, wherein the DAG is composed of a group of nodes having an associated relationship in the DAG;
[0085] The scheduling module 333 is used to determine the execution order of the data conversion tasks obtained by the determination module 332 based on the data conversion task scheduling framework corresponding to the directed acyclic graph.
[0086] Furthermore, the source end and the target end are relational databases, and the source data and the target data are data tables in the relational databases;
[0087] The splitting module 331 is specifically configured to split the source data table according to the preset dimension and the first association relationship, and to construct an index for the split data to obtain the atomic data.
[0088] Furthermore, the determination module 332 is specifically used to determine the dependency relationship between the atomic data according to the cutover rules; construct a directed acyclic graph with the atomic data as nodes and the dependency relationship as edges; and determine a group of atomic data with serial dependency in the directed acyclic graph as a data conversion task.
[0089] Furthermore, the conversion unit 34 is specifically configured to dynamically control the concurrency of the data conversion tasks according to the computing performance of the target computing medium.
[0090] Furthermore, the determining unit 33 is further configured to, when the source data obtained from the source end has a second data format, determine a conversion task corresponding to the source data having the second data format;
[0091] The conversion unit 34 is further configured to execute the conversion task using a preset scheduling program, wherein the scheduling program is configured to cooperate with a data conversion task scheduling framework corresponding to the directed acyclic graph to determine the execution order of the conversion tasks.
[0092] Further, such as Figure 7 As shown, the device also includes:
[0093] The acquisition unit 32 is further configured to acquire a data audit script corresponding to the target computing medium;
[0094] The audit unit 36 is configured to audit the target data and the source data using the data audit script obtained by the acquisition unit 32 before the second transmission unit 35 transmits the target data to the target end.
[0095] Furthermore, the conversion unit 34 is further configured to visualize the execution process of the data conversion task using the directed acyclic graph based on the data conversion task scheduling framework corresponding to the directed acyclic graph.
[0096] Furthermore, an embodiment of the present invention also proposes a data cutover system, which includes a source end, a data cutover end, and a target end, wherein the source end and the target end are both used to store data, the source end stores the old system data to be cutover, and the target end stores the new system data after the cutover.
[0097] The data cutover end is mainly used for target computing media, which is used to convert source data into target data, including but not limited to databases, application systems, big data platforms, etc.
[0098] The data cutover end responds to the cutover request triggered by the user, calls the preset data transmission service, and obtains the source data from the source end through the data transmission service. The cutover request contains the source end identification information corresponding to the data to be obtained and the identification information corresponding to the source data, so that the data transmission service can obtain the source data to be cutover from the specified source end according to the cutover request. After obtaining the source data, the data cutover end will apply the above Figure 1 or Figure 2 The data cutover method shown in converts the source data into target data, and publishes the target data to the target end, completing the data cutover operation on the source data.
[0099] In addition, an embodiment of the present invention further provides a processor, which is used to run a program, wherein the program executes the above Figure 1-2 A data cutover method is provided by any one of the embodiments shown.
[0100] In addition, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the above Figure 1-2 The data cutover method described in any one of the embodiments shown.
[0101] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0102] It is understood that the relevant features of the above methods and devices can be referenced to each other. In addition, the terms "first" and "second" in the above embodiments are used to distinguish between the embodiments, and do not represent the advantages and disadvantages of the embodiments.
[0103] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0104] The algorithm and demonstration provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present invention is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present invention described herein, and the above description of specific languages is for the purpose of disclosing preferred embodiments of the present invention.
[0105] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0106] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0107] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data cutover device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data cutover device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0108] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data switching device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0109] These computer program instructions can also be loaded onto a computer or other programmable data switching device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0110] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0111] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0112] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0113] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0114] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0115] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A data cutover method, applied to a target computing medium, wherein the target computing medium is used to convert source data into target data, the method comprising: Use data transfer services to obtain source data from the source; Obtaining a cutover rule corresponding to the source data and applicable to the target computing medium; Determining a data conversion task having an execution order based on the first association relationship of the source data and the cutover rule, wherein the data conversion task is a conversion operation corresponding to converting the source data having the first association relationship into target data having the second association relationship according to the cutover rule, the data conversion task being composed of a group of nodes having association relationships in a directed acyclic graph, the directed acyclic graph being constructed from atomic data and the cutover rule, the atomic data being obtained by splitting the source data according to a preset splitting rule; Execute the data conversion task according to the execution order to obtain target data having a second association relationship; The target data is transmitted to the target end.
2. The method according to claim 1, characterized in that The method further comprises: An execution order of the data conversion tasks is determined based on a data conversion task scheduling framework corresponding to the directed acyclic graph.
3. The method according to claim 2, characterized in that The source end and the target end are relational databases, and the source data and the target data are data tables in the relational databases; Splitting the source data according to a preset splitting rule includes: The source data table is split according to the preset dimension and the first association relationship, and an index is constructed for the split data to obtain the atomic data.
4. The method according to claim 2, characterized in that Constructing a directed acyclic graph for expressing a data conversion task according to the atomic data and the cutover rule, including: Determining the dependency relationship between the atomic data according to the cutover rule; Constructing a directed acyclic graph with the atomic data as nodes and the dependency relationships as edges; A group of atomic data with a serial dependency relationship in the directed acyclic graph is determined as a data conversion task.
5. The method according to claim 1, wherein Executing the data conversion task according to the execution order includes: The concurrency number of the data conversion tasks is dynamically controlled according to the computing performance of the target computing medium.
6. The method according to claim 2, characterized in that The method further comprises: When the source data obtained from the source end has a second data format, determining a conversion task corresponding to the source data having the second data format; The conversion tasks are executed using a preset scheduling program, and the scheduling program is used to cooperate with a data conversion task scheduling framework corresponding to the directed acyclic graph to determine the execution order of the conversion tasks.
7. The method according to claim 1, characterized in that Before transmitting the target data to the target end, the method further includes: Obtaining a data audit script corresponding to the target computing medium; The data audit script is used to audit the target data and the source data.
8. The method according to claim 2, characterized in that The executing the data conversion task according to the execution order includes: The data conversion task scheduling framework corresponding to the directed acyclic graph is based on which the execution process of the data conversion task is visualized using the directed acyclic graph.
9. The method according to any one of claims 1 to 8, characterized in that The target computing medium includes: database, application system, and big data platform.
10. A data cutover device, applied to a target computing medium, the target computing medium being used to convert source data into target data, the device comprising: A first transmission unit, configured to obtain source data from a source end using a data transmission service; an acquiring unit, configured to acquire a cutover rule corresponding to the source data obtained by the first transmission unit and applicable to the target computing medium; a determining unit, configured to determine a data conversion task having an execution order based on the first association relationship of the source data and the cutover rule obtained by the obtaining unit, wherein the data conversion task is a conversion operation corresponding to converting the source data having the first association relationship into target data having the second association relationship according to the cutover rule, the data conversion task being composed of a group of nodes having association relationships in a directed acyclic graph, the directed acyclic graph being constructed from atomic data and the cutover rule, the atomic data being obtained by splitting the source data according to a preset splitting rule; a conversion unit, configured to execute the data conversion task determined by the determination unit according to the execution order, to obtain target data having a second association relationship; The second transmission unit is configured to transmit the target data obtained by the conversion unit to a target end.
11. A data cutover system, comprising: A source end, a data cutover end, and a target end; wherein the data cutover end is applied to a target computing medium, and the target computing medium is used to convert the source data into the target data; The data cutover end is configured to call a data transmission service according to a cutover request to obtain source data from a source end, convert the source data into target data based on the data cutover method according to any one of claims 1 to 9, and transmit the target data to the target end; The source end is used to transmit the source data corresponding to the cutover request to the target computing medium corresponding to the data cutover end by using a data transmission service.
12. A processor, characterized in that: The processor is configured to run a program, wherein the program, when running, executes the data cutover method according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, wherein when the computer program is run, it controls the device where the computer-readable storage medium is located to execute the data cutover method according to any one of claims 1 to 9.
Citation Information
Patent Citations
A configurable data cleaning system and method
CN108984652A