Data migration method of heterogeneous database based on reinforcement learning

By introducing forward-looking transfer strategy optimization mechanism and cross-domain knowledge transfer of deep reinforcement learning in data migration, the problem of insufficient robustness and adaptability of data migration in the existing technology is solved, and a more efficient and flexible data migration strategy is achieved.

CN120179629AActive Publication Date: 2025-06-20LOGISTICAL ENGINEERING UNIVERSITY OF PLA

Patent Information

Application Number
CN202510249098.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing data transfer method based on reinforcement learning has shortcomings in terms of transfer robustness and adaptability, and is prone to falling into local optimal solutions.

Method used

Adopt the forward-looking transfer strategy optimization mechanism of deep reinforcement learning, combined with cross-domain knowledge transfer, and optimize the data migration process by predicting future needs and environmental changes. Specific steps include extracting data to be migrated from heterogeneous databases, calculating migration path similarity, calculating prospective factors and task complexity factors, updating policy parameters, correcting meta gradients, and initializing the policy network to generate migration policies.

Benefits of technology

Improve the robustness and adaptability of data migration, avoid local optimal solutions, optimize migration strategies to adapt to changing environments and needs, and improve migration quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179629A_ABST
    Figure CN120179629A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data migration, particularly relates to a data migration method of a heterogeneous database based on reinforcement learning, and aims at solving the problems that an existing data migration method based on reinforcement learning is poor in migration robustness and adaptability and prone to falling into a local optimal solution. The method comprises the following steps: extracting data to be migrated and taking a corresponding migration task as a source task; calculating the migration path similarity between the target task and each source task, and selecting a small batch of tasks; calculating a foresight factor and a task complexity factor; calculating a strategy gradient and updating strategy parameters so as to update the strategy network; calculating element gradients, and correcting the element gradients; updating global parameters based on the corrected element gradient; and initializing the policy network based on the global parameters, generating a migration policy through the initialized policy network, and then migrating the source data. According to the method, the robustness and adaptability of data migration are improved, and the problem of a local optimal solution is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data migration, and particularly relates to a data migration method, system, device and computer-readable storage medium for heterogeneous databases based on reinforcement learning. Background Art

[0002] In the information age, data migration is a common task, especially for data migration between heterogeneous databases. Traditional data migration methods often rely on manual configuration and fixed rules, which are not only time-consuming and laborious, error-prone, but also difficult to adapt to the changing environment and requirements, resulting in poor migration effects.

[0003] With the development of big data and artificial intelligence technologies, reinforcement learning (RL), as an effective automatic optimization method, has gradually been applied to data migration tasks. The reinforcement learning method automatically optimizes the migration strategy through interaction with the environment, improves the migration effect, and can adapt to different environments and tasks, with high flexibility. Compared with traditional rule-based methods, reinforcement learning can continuously adjust and optimize the migration strategy through trial-and-error learning, thus showing stronger adaptability in complex and changing environments.

[0004] However, the existing data migration methods based on reinforcement learning still have some deficiencies. First, these methods mainly focus on the optimization of the current task and lack the ability to predict future requirements and environmental changes. For example, during the data migration process, future data volume, data structure changes, business requirement adjustments, etc. may have a significant impact on the migration strategy, while the existing methods often cannot foresee these changes in advance, resulting in insufficient robustness and adaptability of the migration. Second, the training process of the reinforcement learning model usually requires a large amount of data and computing resources, which is a challenge for environments with limited resources. In addition, reinforcement learning methods are prone to falling into local optimal solutions when dealing with large-scale and high-dimensional data migration tasks, affecting the overall migration effect.

[0005] Based on this, the present invention proposes a data migration method for heterogeneous databases based on reinforcement learning, an optimization mechanism for forward-looking migration strategies based on deep reinforcement learning, combined with cross-domain knowledge transfer, to optimize the data migration process by predicting future requirements and environmental changes. Summary of the Invention

[0006] To solve the above problems in the prior art, that is, to solve the problems of poor robustness and adaptability of the existing data migration methods based on reinforcement learning and being prone to falling into local optimal solutions, in the first aspect of the present invention, a data migration method for heterogeneous databases based on reinforcement learning is proposed, and the method includes:

[0007] S10. Extract the data to be migrated from each heterogeneous database as source data, and use the migration tasks corresponding to the source data as source tasks;

[0008] S20. Determine the migration task corresponding to the target database as the target task; calculate the migration path similarity between the target task and each source task and sort them in descending order, and use the source tasks corresponding to the top k migration path similarities before sorting as mini-batch tasks;

[0009] S30. Calculate the forward-looking factor and task complexity factor corresponding to each mini-batch task;

[0010] S40. Combine the forward-looking factor and the task complexity factor to calculate the policy gradient; update the policy parameters based on the policy gradient, and then update the policy network;

[0011] S50. Calculate the meta-gradient according to the updated policy network, and correct the meta-gradient in combination with the task correlation matrix; update the global parameters based on the corrected meta-gradient;

[0012] S60. Initialize the policy network based on the global parameters, generate a migration policy through the initialized policy network, and then migrate the source data.

[0013] In some preferred embodiments, to calculate the migration path similarity between the target task and each source task, the method is as follows:

[0014] Extract the feature vectors of the target task and the source task respectively as the first feature vector and the second feature vector;

[0015] Perform standardization processing on the first feature vector and the second feature vector respectively;

[0016] Calculate the similarity between the standardized first feature vector and the standardized second feature vector through a similarity calculation method as the migration path similarity.

[0017] In some preferred embodiments, the calculation method of the forward-looking factor is as follows:

[0018] where p i represents the forward-looking factor corresponding to the i-th mini-batch task. The forward-looking factor represents the influence factor of the current task on future tasks, t0 represents the task start time point, t f represents the last time point within a future period of time, Demand(t) represents the time migration demand data at the current time point t, which is predicted through a deep learning model based on historical data, and λ represents the decay rate.

[0019] In some preferred embodiments, the task complexity factor is calculated as follows:

[0020] Extract the sub-feature vectors corresponding to the i-th mini-batch task;

[0021] Calculate the weights corresponding to each sub-feature vector, and weight the corresponding sub-feature vectors with the weights; sum the weighted sub-feature vectors to obtain the weighted feature sum of the i-th mini-batch task as the first summation result;

[0022] Calculate the ratio of the first summation result to the second summation result as the task complexity factor; the second summation result is the maximum value of the weighted feature sums of all mini-batch tasks.

[0023] In some preferred embodiments, the policy parameters are updated based on the policy gradient, and the method is as follows:

[0024] where θ′ represents the updated policy parameter, θ represents the unupdated policy parameter, g i represents the policy gradient, S i,target represents the migration path similarity, γ t = e -λt represents the time decay factor, represents the task complexity factor, represents the performance gradient of the i-th mini-batch task, J(■) represents the expected return, π θ represents the policy corresponding to the policy parameter, represents the policy π θ and the trajectory generated by interacting with the environment.

[0025] In some preferred embodiments, the meta-gradient is corrected in combination with the task correlation matrix, and the method is as follows:

[0026] where G′ i represents the corrected meta-gradient, G i represents the uncorrected meta-gradient, C i represents the consistency index value, that is, the ratio of the number of matching data between the migrated data and the source data to the total number of data, I i represents the integrity index value after migration, that is, the ratio of the number of migrated data to the number of source data, V i represents the migration speed index value, that is, the ratio of the amount of migrated data to the time used for migration, U i represents the data usage efficiency index value after migration, that is, the ratio of the average response time in the target database to the benchmark response time, R ij represents the mini-batch task Ti The correlation with the small batch task T j Among them, Q i represents the priority factor of the small batch task T i , Prioroty(T i ) represents the priority of the i-th small batch task, π θ ′ represents the policy updated based on the updated policy parameters represents the policy π θ ′ and the trajectory generated by interacting with the environment

[0027] In some preferred embodiments, the global parameters are updated based on the corrected meta-gradient, and the method is as follows:

[0028] Calculate the arithmetic mean of the corrected meta-gradients corresponding to k small batch tasks

[0029] Sum the unupdated policy parameters and the arithmetic mean as the global parameters

[0030] In the second aspect of the present invention, a data migration system for heterogeneous databases based on reinforcement learning is proposed. The system includes:

[0031] A data extraction module configured to extract the data to be migrated from each heterogeneous database as source data, and use the migration tasks corresponding to each source data as source tasks

[0032] A similarity calculation module configured to determine the migration task corresponding to the target database as the target task; calculate the migration path similarities between the target task and each source task and sort them in descending order, and use the source tasks corresponding to the top k migration path similarities before sorting as small batch tasks

[0033] A factor calculation module configured to calculate the forward-looking factor and the task complexity factor corresponding to each small batch task

[0034] A policy update module that combines the forward-looking factor and the task complexity factor to calculate the policy gradient; updates the policy parameters based on the policy gradient, and further updates the policy network

[0035] A gradient correction module configured to calculate the meta-gradient according to the updated policy network, and correct the meta-gradient in combination with the task correlation matrix; update the global parameters based on the corrected meta-gradient

[0036] A data migration module configured to initialize the policy network based on the global parameters, generate a migration policy through the initialized policy network, and then migrate the source data

[0037] In the third aspect of the present invention, a data migration device for heterogeneous databases based on reinforcement learning is proposed. The device includes:

[0038] At least one processor, and a memory communicatively connected to the at least one processor;

[0039] Wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned data migration method for heterogeneous databases based on reinforcement learning.

[0040] In a fourth aspect of the present invention, a computer-readable storage medium is proposed, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by a computer to implement the above-mentioned data migration method for heterogeneous databases based on reinforcement learning.

[0041] Advantages of the present invention:

[0042] The present invention improves the robustness and adaptability of data migration, and solves the problem of local optimal solutions.

[0043] 1) The present invention introduces a forward-looking factor, predicts future demands and environmental changes through time series analysis and deep learning models, optimizes the migration strategy, and improves the migration quality and efficiency; utilizes cross-domain knowledge transfer technology to apply knowledge and experience in other fields to the current task, improving the generalization ability and adaptability of the migration strategy; combines a multi-level feedback mechanism, comprehensively considers the similarity of tasks, time decay, forward-looking factor, task complexity factor, and task priority factor, optimizes the policy gradient, and improves the accuracy and effectiveness of the migration strategy; through the task complexity factor and task priority factor, more accurately evaluates the difficulty and importance of tasks, optimizes the selection and scheduling of tasks, improves the migration efficiency, and thus solves the problem that traditional data migration methods are difficult to adapt to the changing environment and demands, resulting in poor migration effects and being prone to falling into local optimal solutions.

[0044] 2) The present invention automatically optimizes the migration strategy, reduces manual configuration and intervention, and improves work efficiency. Description of the Drawings

[0045] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious.

[0046] Figure 1 It is a schematic flowchart of a data migration method for a heterogeneous database based on reinforcement learning according to an embodiment of the present invention. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] The following further elaborates on the present application with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and do not limit the invention. Additionally, it should be noted that for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0049] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0050] A data migration method for heterogeneous databases based on reinforcement learning according to the first embodiment of the present invention is as Figure 1 shown and includes the following steps:

[0051] S10: Extract the data to be migrated from each heterogeneous database as source data, and use the migration tasks corresponding to each source data as source tasks;

[0052] S20: Determine the migration task corresponding to the target database as the target task; calculate the migration path similarity between the target task and each source task and sort them in descending order, and use the source tasks corresponding to the top k migration path similarities before sorting as mini-batch tasks;

[0053] S30: Calculate the forward-looking factor and task complexity factor corresponding to each mini-batch task;

[0054] S40: Combine the forward-looking factor and the task complexity factor to calculate the policy gradient; update the policy parameters based on the policy gradient, and then update the policy network;

[0055] S50: Calculate the meta-gradient according to the updated policy network, and correct the meta-gradient in combination with the task correlation matrix; update the global parameters based on the corrected meta-gradient;

[0056] S60: Initialize the policy network based on the global parameters, generate a migration policy through the initialized policy network, and then migrate the source data.

[0057] To more clearly illustrate a data migration method for heterogeneous databases based on reinforcement learning of the present invention, each step in an embodiment of the method of the present invention will be elaborated in detail below with reference to the accompanying drawings.

[0058] In the present invention, when faced with multiple similar but not exactly identical migration tasks, cross-domain knowledge transfer technology is introduced, allowing for the learning of general knowledge and strategies from completed migration tasks and applying them to new migration tasks, thereby quickly adapting to new tasks, accelerating the learning speed of new tasks, and also increasing the success rate of migration.

[0059] The application of cross-domain knowledge transfer in reinforcement learning can be achieved in various ways, including transfer learning, meta-learning (meta-learning, also known as "learning how to learn", refers to the ability of a model to quickly adapt to new tasks by learning from multiple tasks), etc.

[0060] The first embodiment of the present invention provides a data migration method for heterogeneous databases based on reinforcement learning, adopting a meta-learning method (specifically: each data migration task is regarded as a task, and the goal of cross-domain knowledge transfer is to learn from previous tasks to better and more quickly complete new tasks), and optimizing the migration strategy and improving the migration quality and efficiency by introducing a forward-looking factor, a multi-level feedback mechanism, a task complexity factor, and a task priority factor. Specifically as follows:

[0061] S10, Extract the data to be migrated from each heterogeneous database as source data, and use the migration tasks corresponding to each source data as source tasks;

[0062] In this embodiment, first extract the data that needs to be migrated from existing heterogeneous databases. Taking the material and equipment management system as an example, the source data includes personnel data, financial data, and material data. Extract information such as material name, quantity, storage location, and supply unit from the material management system; extract information such as personnel name, position, affiliated unit, and contact information from the personnel management system; extract information such as budget, expenditure, income, and audit records from the financial management system.

[0063] According to the type and characteristics of the source data, define the source tasks and clarify the scope of the source tasks to facilitate subsequent task selection and strategy optimization. For example, T1: Personnel data migration, T2: Financial data migration;

[0064] S20, Determine the migration task corresponding to the target database as the target task; calculate the migration path similarity between the target task and each source task and sort them in descending order, and use the source tasks corresponding to the top k migration path similarities before sorting as mini-batch tasks;

[0065] The goal of data migration is to integrate the data in multiple heterogeneous databases into a unified target database to improve the efficiency of data management and use. To ensure the success of data migration, it is necessary to clearly define the target task according to the requirements of the target database.

[0066] In this embodiment, the structure, functions, and business requirements of the target database are understood. Structure requirements: Understand the table structure, field definitions, primary key and foreign key relationships, etc. of the target database; Function requirements: Understand the functions that the target database needs to support, such as data query, report generation, data analysis, etc.; Business requirements: Understand the business processes that the target database needs to support, such as material management, personnel management, financial management, etc.

[0067] Decompose the requirements of the target database into specific migration tasks, that is, target tasks. Each task corresponds to a specific data migration operation, and determine the mapping relationship between the source data and the target data to ensure the consistency and integrity of the data during the migration process. For example, migrate the data in the personnel management system to the personnel management table in the target database, and map fields such as "personnel name", "position", "affiliated unit", and "contact information" in the personnel management system to the corresponding fields in the target database; migrate the data in the financial management system to the financial management table in the target database, and map fields such as "budget", "expenditure", "income", and "audit record" in the financial management system to the corresponding fields in the target database.

[0068] Then, initialize the parameters θ of the policy network and the parameters of the value network; select the top k source tasks with the highest similarity of the migration path to the target task from the source task set as the mini-batch tasks.

[0069] Among them, the calculation method of the migration path similarity is as follows:

[0070] Extract the feature vectors of the target task and the source task respectively as the first feature vector and the second feature vector; perform standardization processing on the first feature vector and the second feature vector respectively;

[0071] Calculate the similarity between the standardized first feature vector and the standardized second feature vector through the similarity calculation method as the migration path similarity; the similarity calculation method preferably uses the cosine similarity formula to calculate in the present invention, and can also preferably use a neural network model for calculation, such as the LSTM model. In the present invention, the LSTM model is constructed based on the sequentially connected input layer, convolutional layer (for feature extraction), Transformer encoding layer, gated recurrent unit layer, fully connected layer, and output layer. Among them, the output layer of the Transformer encoding layer is connected with the output residual of the gated recurrent unit layer as the input of the fully connected layer; the Transformer encoding layer is used to perform dynamic position encoding on the output feature vector to obtain the encoded vector; respectively through the global attention mechanism (that is, the global vector Attention = X + α2.Attention(Q, K, V), α2 = α 1+Linear(C), α1 = LayerNorm(ReLU(Linear(X))) + X, C = Condition(X), where X represents the input feature vector, Linear represents the linear transformation, LayerNorm represents layer normalization, ReLU represents the activation function, Attention(Q, K, V) is the standard self-attention mechanism, and Condition represents the conditional variable). The global vector and the local vector are obtained by weighting the encoded vector using the local attention mechanism (the local attention mechanism applied in the present invention is an existing technology and will not be elaborated here), and the global vector and the local vector are fused as the output.

[0072] S30, calculate the look-ahead factor and the task complexity factor corresponding to each mini-batch task;

[0073] In heterogeneous database data migration, future requirements and environments may change. Therefore, the migration strategy needs to have a certain degree of forward-lookingness and be able to predict and adapt to future changes.

[0074] In this embodiment, a forward-looking migration strategy optimization mechanism based on deep reinforcement learning is introduced. By analyzing time series and using a deep learning model to predict future requirements, the adaptability and robustness of the migration strategy are improved.

[0075] The look-ahead factor (i.e., the impact factor of the current task on future tasks) is specifically as follows:

[0076] where p i represents the look-ahead factor corresponding to the i-th mini-batch task. The look-ahead factor represents the impact factor of the current task on future tasks. t0 represents the task start time point, and t - t0 represents the time difference, which is used to capture the demand changes of the task in a future period. f t represents the last time point in a future period. Demand(t) represents the time migration demand data at the current time point t, which is predicted based on historical data (i.e., the migration task records in the past period, including migration time, source data, target data, migration strategy, migration results, etc.) through a deep learning model (an existing technology and will not be elaborated here), and λ represents the decay rate.

[0077] Through the time difference, the look-ahead factor can capture the demand changes of the task in a future period, thus better reflecting the impact of the current task on future tasks. This method is particularly useful in data migration because it can help optimize the migration strategy and ensure that the migration process can adapt to future business requirements. In data management, this calculation method of the look-ahead factor can significantly improve the effect and efficiency of data migration.

[0078] The task complexity factor needs to be able to accurately evaluate the complexity of the task, taking into account various factors such as data scale, data type, migration path, etc. The present invention uses a multi-factor comprehensive evaluation model, combining deep learning and expert knowledge, to calculate the task complexity factor, specifically as follows:

[0079] Each mini-batch task T i has multiple features f (i.e., sub-feature vectors), and these features can include data scale, data type diversity, migration path length, data dependence, migration frequency, etc. In the present invention, first, extract the sub-feature vectors corresponding to the i-th mini-batch task; calculate the weights corresponding to the sub-feature vectors, and weight the corresponding sub-feature vectors by the weights; sum the weighted sub-feature vectors to obtain the weighted feature sum of the i-th mini-batch task as the first summation result; calculate the ratio of the first summation result to the second summation result as the task complexity factor; the second summation result is the maximum value of the weighted feature sums of all mini-batch tasks. As shown in the following formula:

[0080] Among them, represents the task complexity factor, ∑ f∈F ω f ·Feature f (T i ) represents the weighted feature sum of the mini-batch task, F represents the feature set, ω f represents the weight of the sub-feature vector f, Feature f (T i ) represents the sub-feature vectors corresponding to the mini-batch task T i , that is, the values of the mini-batch task on the sub-feature vectors, T j represents the set of all mini-batch tasks, and max represents taking the maximum value of the weighted feature sums of all tasks.

[0081] S40. Combine the foresight factor and the task complexity factor to calculate the policy gradient; update the policy parameters based on the policy gradient, and then update the policy network;

[0082] In this embodiment, for each mini-batch task, use the policy to interact with the environment to collect a certain number of trajectories; then calculate the foresight factor and the task complexity factor;

[0083] Combine the foresight factor and the task complexity factor to calculate the policy gradient, and update the policy parameters based on the policy gradient. The specific process is as follows:

[0084] Among them, θ′ represents the updated policy parameter, θ represents the unupdated policy parameter, gi Denote the policy gradient as S i,target Denote the migration path similarity as γ t = e -λt Denote the time decay factor (During the data migration process, the early migration paths and policies may not be as effective as the later paths and policies. Therefore, a time decay factor can be introduced to reduce the impact of early tasks), Denote the task complexity factor, Denote the performance gradient of the i-th mini-batch task, J(■) denotes the expected return, and π θ Denote the policy corresponding to the policy parameters, Denote the policy π θ The trajectory generated by interacting with the environment.

[0085] In reinforcement learning, the performance gradient is used to guide the update of policy parameters to maximize the expected return. The performance gradient reflects the sensitivity of policy parameters to the expected return. In the present invention, the calculation process of the performance gradient is as follows:

[0086] When using the policy π θ to interact with the environment, one or more trajectories are generated (the trajectory includes the state s, the action a, and the reward r). For each time step in each trajectory, calculate the action a t in the state s t The logarithmic probability logπ θ (a t |s t ).

[0087] For each time step t in each trajectory, calculate the cumulative reward from t to the end of the trajectory: γ is the discount factor, which is set according to the actual situation; for each time step in each trajectory, calculate the gradient term Add up the gradient terms of all time steps to obtain the performance gradient of this trajectory. If there are multiple trajectories, then average the gradients of all trajectories as the final performance gradient.

[0088] S50. According to the updated policy network, calculate the meta-gradient, and correct the meta-gradient in combination with the task correlation matrix; update the global parameters based on the corrected meta-gradient;

[0089] In this embodiment, the updated policy is used to interact with the environment again on small-batch tasks to collect new trajectories, and then calculate the introduction of a multi-level feedback mechanism (a single migration effect feedback may not be sufficient to comprehensively evaluate the quality of migration), including the consistency index value of the data after migration (i.e., the ratio of the number of data items in the migrated data that match the source data to the total number of data items, that is, by comparing each item of the migrated data with the source data item by item to obtain the number of matching data items, and taking the ratio of the number of matching data items to the total number of data items (i.e., the amount of migrated data) as the consistency index value), the integrity index value (i.e., the ratio of the number of data items after migration to the number of data items in the source data), the migration speed index value (i.e., the ratio of the amount of migrated data to the time taken for migration), and the data usage efficiency index value after migration (i.e., the ratio of the average response time in the target database to the baseline response time (set according to specific circumstances)).

[0090] Use attention to calculate the task correlation matrix, and then calculate the meta-gradient, and correct the meta-gradient by combining the multi-level feedback and the task correlation matrix; update the global parameters based on the corrected meta-gradient, and the specific process is as follows:

[0091] Among them, G′ i represents the corrected meta-gradient, G i represents the uncorrected meta-gradient, C i represents the consistency index value, that is, the ratio of the number of data items in the migrated data that match the source data to the total number of data items, I i represents the integrity index value after migration, that is, the ratio of the number of data items after migration to the number of data items in the source data, V i represents the migration speed index value, that is, the ratio of the amount of migrated data to the time taken for migration, U i represents the data usage efficiency index value after migration, that is, the ratio of the average response time in the target database to the baseline response time, R ij represents the small-batch task T i and the small-batch task T j The correlation between them (the correlation between different tasks can further affect the calculation of the meta-gradient. By introducing the task correlation matrix, the meta-gradient can be adjusted more precisely), Q i represents the small-batch task T i The priority factor of, Prioroty(T i ) represents the priority of the i-th small-batch task, π θ ′ represents the policy updated based on the updated policy parameters, represents the policy π θ ′ and the trajectory generated by interacting with the environment.

[0092] Update the global parameters based on the corrected meta-gradient, and the method is as follows:

[0093] Calculate the arithmetic mean of the corrected meta-gradients corresponding to k mini-batch tasks; sum the unupdated policy parameters and the arithmetic mean as the global parameters.

[0094] S60, Initialize the policy network based on the global parameters, generate a migration policy through the initialized policy network, and then migrate the source data.

[0095] In this embodiment, use the global parameters updated by meta-update to initialize the policy network, generate a migration policy (including data extraction, transformation (i.e., according to the generated migration policy, perform necessary transformation on the extracted data to ensure that the data format meets the requirements of the target database) and loading), and perform data migration in the actual environment. Or after initializing the policy network, extract a part of the data from the target database as a fine-tuning data set, generate a migration policy for data migration, and fine-tune the policy network according to the results and feedback information of the data migration (such as consistency, integrity), and then perform data migration in the actual environment after fine-tuning.

[0096] In summary, the present invention improves the performance in heterogeneous database data migration by introducing a forward-looking factor, a multi-level feedback mechanism, a task complexity factor, and a task priority factor, enabling the model to more flexibly adapt to task changes and exhibit higher migration quality and efficiency on new tasks.

[0097] A heterogeneous database data migration system based on reinforcement learning according to the second embodiment of the present invention, the system includes:

[0098] A data extraction module configured to extract the data to be migrated from each heterogeneous database as source data, and use the migration tasks corresponding to each source data as source tasks;

[0099] A similarity calculation module configured to determine the migration task corresponding to the target database as the target task; calculate the migration path similarities between the target task and each source task and sort them in descending order, and use the source tasks corresponding to the top k migration path similarities before sorting as mini-batch tasks;

[0100] A factor calculation module configured to calculate the forward-looking factor and the task complexity factor corresponding to each mini-batch task;

[0101] A policy update module, which combines the forward-looking factor and the task complexity factor to calculate the policy gradient; updates the policy parameters based on the policy gradient, and then updates the policy network;

[0102] The gradient correction module is configured to calculate the meta-gradient according to the updated policy network, and correct the meta-gradient in combination with the task correlation matrix; and update the global parameters based on the corrected meta-gradient.

[0103] The data migration module is configured to initialize the policy network based on the global parameters, generate a migration policy through the initialized policy network, and then migrate the source data.

[0104] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related explanations of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0105] It should be noted that the data migration system for heterogeneous databases based on reinforcement learning provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as an improper limitation of the present invention.

[0106] A data migration device for heterogeneous databases based on reinforcement learning according to the third embodiment of the present invention includes at least one processor; and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned data migration method for heterogeneous databases based on reinforcement learning.

[0107] A computer-readable storage medium according to the fourth embodiment of the present invention stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned data migration method for heterogeneous databases based on reinforcement learning.

[0108] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related explanations of the above-described data migration device for heterogeneous databases based on reinforcement learning and the computer-readable storage medium can refer to the corresponding process in the foregoing method examples, and will not be repeated here.

[0109] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0110] The terms "first", "second", "third", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.

[0111] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A data migration method for heterogeneous databases based on reinforcement learning, characterized in that: The method includes: S10, extracting data to be migrated from each heterogeneous database as source data, and taking the migration task corresponding to each source data as the source task; S20, determining the migration task corresponding to the target database as the target task; calculating the similarity between the migration paths of the target task and each source task and sorting them in descending order, and taking the source tasks corresponding to the top k migration path similarities as small batch tasks; S30, calculating the forward-looking factor and task complexity factor corresponding to each small batch task; S40, calculating a policy gradient by combining the forward-looking factor and the task complexity factor; updating a policy parameter based on the policy gradient, and then updating a policy network; S50, calculating a meta-gradient according to the updated policy network, and correcting the meta-gradient in combination with the task relevance matrix; and updating the global parameters based on the corrected meta-gradient; S60, initializing a policy network based on the global parameters, generating a migration strategy through the initialized policy network, and then migrating the source data.

2. The data migration method of heterogeneous databases based on reinforcement learning according to claim 1 is characterized in that: The similarity of the migration paths between the target task and each source task is calculated as follows: Extracting feature vectors of the target task and the source task respectively as a first feature vector and a second feature vector; performing standardization processing on the first eigenvector and the second eigenvector respectively; The similarity between the first eigenvector after the normalization process and the second eigenvector after the normalization process is calculated by a similarity calculation method as the migration path similarity.

3. The data migration method of heterogeneous database based on reinforcement learning according to claim 2 is characterized in that: The forward-looking factor is calculated as follows: Among them, p i represents the forward-looking factor corresponding to the i-th small batch task. The forward-looking factor represents the impact factor of the current task on future tasks. t0 represents the starting time of the task, and t f It represents the last time point in the future. Demand(t) represents the time migration demand data at the current time point t, which is predicted by the deep learning model based on historical data. λ represents the decay rate.

4. The data migration method for heterogeneous databases based on reinforcement learning according to claim 3 is characterized in that: The task complexity factor is calculated as follows: Extract each sub-feature vector corresponding to the i-th mini-batch task; Calculate the weight corresponding to each sub-feature vector, and weight the corresponding sub-feature vector by the weight; sum the weighted sub-feature vectors to obtain the weighted feature sum of the i-th small batch task as the first summation result; The ratio of the first summation result to the second summation result is calculated as the task complexity factor; the second summation result is the maximum value of the weighted feature sums in all small batch tasks.

5. The data migration method of heterogeneous database based on reinforcement learning according to claim 4 is characterized in that: The method for updating the policy parameters based on the policy gradient is as follows: Among them, θ′ represents the updated policy parameters, θ represents the unupdated policy parameters, and g i represents the policy gradient, S i,target represents the similarity of migration paths, γ t =e -λt represents the time decay factor, represents the task complexity factor, represents the performance gradient of the i-th mini-batch task, J(■) represents the expected return, π θ Indicates the strategy corresponding to the strategy parameter. Represents strategy π θ Trajectories generated by interacting with the environment.

6. The data migration method for heterogeneous databases based on reinforcement learning according to claim 5 is characterized in that: The meta-gradient is modified in combination with the task relevance matrix as follows: Among them, G′ i represents the corrected meta-gradient, G i represents the uncorrected meta-gradient, C i Represents the consistency index value, that is, the ratio of the number of data items that match the migrated data and the source data to the total number of data items. i Represents the integrity index value after migration, that is, the ratio of the number of data items after migration to the number of source data items, V i Indicates the migration speed index value, that is, the ratio of the amount of data migrated to the time taken for migration. i represents the data usage efficiency index value after migration, that is, the ratio of the average response time in the target database to the benchmark response time, R ij Represents a small batch task T i With a small batch task T j The correlation between i Represents a small batch task T i Priority factor, Prioroty(T i ) represents the priority of the i-th mini-batch task, π θ′ represents the updated policy based on the updated policy parameters, Represents strategy π θ′ Trajectories generated by interacting with the environment.

7. The data migration method of heterogeneous databases based on reinforcement learning according to claim 6 is characterized in that: The global parameters are updated based on the corrected meta-gradients as follows: Calculate the arithmetic mean of the corrected meta-gradients corresponding to k mini-batches; The unupdated policy parameters are summed with the arithmetic mean to obtain the global parameters.

8. A data migration system for heterogeneous databases based on reinforcement learning, characterized in that: The system comprises: A data extraction module is configured to extract data to be migrated from various heterogeneous databases as source data, and to use migration tasks corresponding to various source data as source tasks; A similarity calculation module is configured to determine the migration task corresponding to the target database as the target task; calculate the similarity between the migration paths of the target task and each source task and sort them in descending order, and take the source tasks corresponding to the top k migration path similarities as small batch tasks; A factor calculation module, configured to calculate a forward-looking factor and a task complexity factor corresponding to each small batch task; A policy updating module, combining the forward-looking factor and the task complexity factor, calculates a policy gradient; updates policy parameters based on the policy gradient, and then updates the policy network; A gradient correction module is configured to calculate a meta-gradient according to the updated policy network, and to correct the meta-gradient in combination with the task relevance matrix; and to update the global parameters based on the corrected meta-gradient; The data migration module is configured to initialize the policy network based on the global parameters, generate a migration strategy through the initialized policy network, and then migrate the source data.

9. An electronic data migration device for heterogeneous databases based on reinforcement learning, characterized in that: The electronic device comprises: at least one processor, and a memory communicatively coupled to at least one of the processors; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the data migration method of heterogeneous databases based on reinforcement learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by a computer to implement the data migration method of a heterogeneous database based on reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge reasoning method based on deep migration reinforcement learning

    CN118734968A

  • Migrating Data In And Out Of Cloud Environments

    US20220019367A1

Cited By

  • Data migration method and device, equipment and storage medium

    CN121092521A