Data synchronization method of heterogeneous database based on reinforcement learning
By introducing an adaptive exploration strategy based on the dynamic characteristics of the environment into the data synchronization method, dynamically adjusting the weight of the objective function, the problem that is difficult to adapt in the dynamic environment in the existing technology is solved, and efficient and accurate data synchronization is achieved.
Patent Information
- Application Number
- CN202510249128.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing data synchronization methods are difficult to adapt in dynamic environments, resulting in synchronization delays or data inconsistencies, and it is impossible to achieve efficient and accurate data synchronization.
A data synchronization method for heterogeneous databases based on reinforcement learning is proposed, and an adaptive exploration strategy based on the dynamic characteristics of the environment is introduced. By dynamically adjusting the weight of the objective function, it can adapt to different environmental changes.
It realizes efficient and accurate data synchronization in a dynamic environment, ensuring that the data in each database remains consistent, and is suitable for synchronization tasks of multiple heterogeneous databases.
Smart Images

Figure CN120179665A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data synchronization, and particularly relates to a data synchronization method, system, electronic device and computer-readable storage medium for heterogeneous databases based on reinforcement learning. Background Art
[0002] In modern information systems, data synchronization of heterogeneous databases is a crucial issue. Traditional data synchronization methods often rely on static configurations and are difficult to adapt to the dynamic changes of data sources. These methods perform well in static environments, but in dynamic environments, such as changes in network conditions and fluctuations in data update frequencies, fixed rules and preset strategies often cannot be adjusted in a timely manner, resulting in synchronization delays or data inconsistencies. For example, in a material and equipment management system, the update frequencies of financial data and personnel data are different, and fixed-rule synchronization may lead to frequent synchronization of financial data while insufficient synchronization of personnel data, or vice versa. In addition, preset synchronization strategies may not be able to handle sudden emergencies, resulting in untimely data synchronization and affecting decision-making efficiency.
[0003] In recent years, reinforcement learning, as a machine learning method, has shown great potential in data synchronization by interacting with the environment to learn optimal strategies. However, existing reinforcement learning methods usually focus on a single goal, such as maximizing synchronization speed or minimizing resource consumption. In practical applications, data synchronization often involves multiple goals, such as synchronization speed, resource utilization rate, and data consistency. Existing multi-objective reinforcement learning methods usually use fixed weights to balance different goals, and these weights may no longer be applicable in different environments, resulting in performance degradation. In addition, existing reinforcement learning methods lack adaptability when dealing with dynamic environmental changes and cannot dynamically adjust the weights of the objective function according to the changes in the current environment, thus unable to achieve efficient and accurate data synchronization.
[0004] Based on this, the present invention proposes a data synchronization method for heterogeneous databases based on reinforcement learning, introducing an adaptive exploration strategy based on the dynamic characteristics of the environment, enabling the algorithm to dynamically adjust the weights of the objective function to adapt to different environmental changes. Summary of the Invention
[0005] In order to solve the above problems in the prior art, that is, to solve the problem that existing data synchronization methods are difficult to adapt in dynamic environments, resulting in synchronization delays or data inconsistencies and unable to achieve efficient and accurate data synchronization, in the first aspect of the present invention, a data synchronization method for heterogeneous databases based on reinforcement learning is proposed. The method includes:
[0006] S10, obtaining the status information of the data to be synchronized currently as the current state;
[0007] S20. Calculate the arithmetic mean of the difference between the action-value function value and the average action-value function value in the current state based on the current state and time step as the confidence level. Combine the confidence level and the current state to calculate the dynamic exploration probability: ∈(s t ,t) = ∈0 × exp(-γ × (confidence(s t ,t) + α d × f d (s t ) + α n × f n (s t ));
[0008] Among them, ∈(s t ,t) represents the dynamic exploration probability, s t represents the current state, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the decay coefficient, confidence(s t ,t) represents the confidence level, f d (s t ) represents the update frequency of the data in the database within a set time unit, f n (s t ) represents the network fluctuation, that is, the stability of the network connection within a set time unit, α d 、a n both represent weight parameters;
[0009] S30. Combine the dynamic exploration probability to select the action of the data to be synchronized currently;
[0010] S40. Execute the selected action, calculate the immediate reward and observe the state of the next time step; Put the current state, the executed selected action, the immediate reward, the state of the next time step, and the importance of the experience into the experience replay pool; Collect samples from the experience replay pool according to the importance of the experience, and update the parameters of the Q network and the target Q network;
[0011] S50. Loop and execute S10 - S40 until the data synchronization reaches the expected performance index, then synchronize the data to be synchronized through the trained Q network.
[0012] In some preferred embodiments, the method of combining the dynamic exploration probability to select the action of the data to be synchronized currently is as follows:
[0013] Generate a random number between 0 and 1;
[0014] If the random number is less than the difference between 1 and the dynamic exploration probability, randomly select an action from the action set as the action for the currently to-be-synchronized data; otherwise, select the action with the highest Q-value in the current state as the action for the currently to-be-synchronized data.
[0015] In some preferred embodiments, the method for obtaining the action with the highest Q-value in the current state is as follows:
[0016] where α t represents the selected action, Q(s t , a; θ) represents the expected return of taking action a in s t , θ represents the parameters of the action value network, random_score(a) is the random score, cur(a) is the curiosity reward score, h(a) represents the historical score, and β1, β2, β3, β4 are all weights.
[0017] In some preferred embodiments, the method for calculating the immediate reward is as follows:
[0018] where R represents the immediate reward, I(t) represents the number of data items recognized at time step t, I max represents the maximum number of recognized data items, T(t) represents the amount of data successfully transmitted at time step t, T max is the maximum amount of transmitted data, ρ(s t , a t ) represents the prediction error, that is, the error between the predicted value of the reward after the agent executes a t in state s t and the actual reward, d(s t , a t ) represents the similarity between the action a t executed by the agent in state s t and the expert demonstration, ω 识别 , ω 传输 both represent weight parameters, δ, k, λ, μ all represent decay coefficients, η, ψ1, ψ2 are all positive constants used to adjust the magnitude of the corresponding rewards.
[0019] In some preferred embodiments, the importance of the experience is calculated based on the TD error of the experience.
[0020] In some preferred embodiments, the method for updating the parameters of the action value network includes the gradient descent method.
[0021] In the second aspect of the present invention, a data synchronization system for heterogeneous databases based on reinforcement learning is proposed. The system includes:
[0022] A status acquisition module, configured to acquire the status information of the data to be synchronized currently as the current status;
[0023] A probability calculation module, configured to calculate the arithmetic mean of the difference between the action value function value and the average action value function value in the current status based on the current status and the time step as the confidence level; combine the confidence level and the current status to calculate the dynamic exploration probability; ∈(s t , t) = ∈0 × exp(-γ × (confidence(s t , t) + α d × f d (s t ) + α n × f n (s t ));
[0024] Wherein, ∈(s t , t) represents the dynamic exploration probability, s t represents the current status, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the attenuation coefficient, confidence(s t , t) represents the confidence level, f d (s t ) represents the update frequency of the data in the database within a set time unit, f n (s t ) represents the network fluctuation, that is, the stability of the network connection within a set time unit, α d and α n both represent weight parameters;
[0025] An action selection module, configured to select the action of the data to be synchronized currently in combination with the dynamic exploration probability;
[0026] A network update module, configured to execute the selected action, calculate the immediate reward and observe the status of the next time step; put the current status, the executed selected action, the immediate reward, the status of the next time step, and the importance of the experience into the experience replay pool; collect samples from the experience replay pool according to the importance of the experience, and update the parameters of the Q network and the target Q network;
[0027] A data synchronization module, configured to loop and execute the status acquisition module - the network update module until the data synchronization reaches the expected performance index, and then synchronize the data to be synchronized through the trained Q network.
[0028] In the third aspect of the present invention, an electronic device is proposed, and the device includes:
[0029] At least one processor, and a memory communicatively connected to the at least one processor;
[0030] Wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned data synchronization method for heterogeneous databases based on reinforcement learning.
[0031] In a fourth aspect of the present invention, a computer-readable storage medium is proposed. The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by a computer to implement the above-mentioned data synchronization method for heterogeneous databases based on reinforcement learning.
[0032] Advantages of the present invention:
[0033] The present invention introduces an adaptive exploration strategy based on the dynamic characteristics of the environment to adapt to different environmental changes, solves the problems of synchronization delay or data inconsistency, and realizes efficient and accurate data synchronization.
[0034] 1) By introducing an adaptive exploration strategy based on the dynamic characteristics of the environment, the present invention not only considers the confidence of the current state, but also combines the dynamic characteristics of the environment, such as data update frequency and network fluctuations, and can dynamically adjust the weights of the objective function to adapt to different environmental changes, so as to achieve efficient and accurate data synchronization and ensure that the data in each database is consistent.
[0035] 2) By introducing reward shaping, curiosity-driven learning, and demonstration learning, the present invention solves the problems of lack of immediate feedback in rewards, low search efficiency, and credit assignment in existing reinforcement learning;
[0036] 3) The present invention is applicable to the synchronization tasks of various heterogeneous databases and has wide applicability and scalability. Description of the Drawings
[0037] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious.
[0038] Figure 1 It is a schematic flowchart of a data synchronization method for heterogeneous databases based on reinforcement learning according to an embodiment of the present invention. Detailed Embodiments
[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] The following further elaborates on the present application with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings.
[0041] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0042] A data synchronization method for heterogeneous databases based on reinforcement learning in the first embodiment of the present invention, as Figure 1 shown, includes the following steps:
[0043] S10. Obtain the status information of the data to be synchronized currently as the current status;
[0044] S20. Based on the current status and time step, calculate the arithmetic mean of the difference between the action value function value and the average action value function value in the current status as the confidence level; combine the confidence level and the current status to calculate the dynamic exploration probability: ∈(s t , t) = ∈0 × exp(-γ × (confidence(s t , t) + α d × f d (s t ) + α n × f n (s t ));
[0045] Among them, ∈(s t , t) represents the dynamic exploration probability, s t represents the current status, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the attenuation coefficient, confidence(s t , t) represents the confidence level, f d (s t ) represents the update frequency of the data in the database within a set time unit, f n (s t ) represents the network fluctuation, that is, the stability of the network connection within a set time unit, α d and αn Both represent weight parameters;
[0046] S30. Combine the dynamic exploration probability to select the action of the data to be synchronized currently;
[0047] S40. Execute the selected action, calculate the immediate reward and observe the state of the next time step; put the current state, the executed selected action, the immediate reward, the state of the next time step, and the importance of the experience into the experience replay pool; sample from the experience replay pool according to the importance of the experience, and update the parameters of the Q network and the target Q network;
[0048] S50. Loop to execute S10 - S40 until the data synchronization reaches the expected performance index, and then synchronize the data to be synchronized through the trained Q network.
[0049] To more clearly illustrate a data synchronization method for heterogeneous databases based on reinforcement learning of the present invention, the following will elaborate on each step in an embodiment of the method of the present invention in conjunction with the accompanying drawings.
[0050] The first embodiment of the present invention provides a data synchronization method for heterogeneous databases based on reinforcement learning. The data synchronization is performed based on the ASRL (Adaptive State Representation Learning) algorithm. By introducing an adaptive exploration strategy based on the dynamic characteristics of the environment, the algorithm can dynamically adjust the weights of the objective function to adapt to different environmental changes. The ASRL algorithm not only considers the confidence of the current state but also combines the dynamic characteristics of the environment, such as data update frequency and network fluctuations, to more finely adjust the exploration probability, thereby achieving efficient and accurate data synchronization and ensuring that the data in each database is consistent. In addition, the present invention solves the problems of lack of immediate feedback of rewards, low search efficiency, credit assignment problems, and local optimal solutions in existing reinforcement learning by introducing reward shaping, curiosity-driven learning, hierarchical reinforcement learning, and demonstration learning. Specifically as follows:
[0051] S10. Obtain the state information of the data to be synchronized currently as the current state;
[0052] In this embodiment, first, the Q network and the target Q network in reinforcement learning are initialized, including initializing the experience replay pool, the initial state, initializing the dynamic exploration probability parameter, the decay coefficient, the data update frequency weight, the network fluctuation weight, etc.
[0053] Then perform state perception to obtain the current state. The state is a description of the environment, representing the current environmental situation. In the scenario of data synchronization, the state includes the state of financial data, the state of personnel data, the network connection state, etc. For example, in a material and equipment management system, the current state includes financial data (such as budget, expenditure, income) and personnel data (such as basic information, training records, health status). For financial data: budget of 1 million, expenditure of 500,000, income of 300,000; for personnel data: total of 100 people, 95 in good health, 5 with injuries or illnesses.
[0054] S20. Calculate the arithmetic mean of the difference between the action value function value and the average action value function value in the current state based on the current state and the time step as the confidence level; combine the confidence level and the current state to calculate the dynamic exploration probability;
[0055] Traditional multi-objective optimization methods usually use fixed weights to balance different objectives. In this embodiment, an adaptive exploration strategy based on the dynamic characteristics of the environment is introduced. This strategy not only considers the confidence level of the current state but also combines the dynamic characteristics of the environment, such as data update frequency and network fluctuations, to more finely adjust the exploration probability, enabling the algorithm to dynamically adjust the weights of the objective function to adapt to different environmental changes. Specifically as follows:
[0056] Calculate the arithmetic mean of the difference between the Q value in the current state (i.e., the Q value of a certain action in the current state. In reinforcement learning, the Q value is usually estimated through a Q network (a neural network). The action value function value is the Q value, representing the expected value of the cumulative reward from the current time step to all future time steps after taking an action in the state. The higher the Q value, the better the long-term benefit of taking this action. Actions include synchronizing financial data, synchronizing personnel data, and pausing synchronization; input the current state and each action value into the Q network to obtain the Q value) and the average Q value (i.e., the average Q value of all actions in the current state) as the confidence level (the smaller the deviation, the higher the confidence level, indicating that the algorithm is more certain about the current state); combine the confidence level and the current state to calculate the dynamically adjusted exploration probability, as shown in the following formula: ∈(s t ,t)=∈0×exp(-γ×(confidence(s t ,t)+α d ×f d (s t )+α n ×f n (s t ));
[0057] Among them, ∈(s t ,t) represents the dynamic exploration probability, s tDenote the current state as \(s\), \(t\) represents the time step, \(\epsilon_0\) represents the initial dynamic exploration probability, which is the exploration probability set at the beginning of the algorithm and is usually set to a relatively high value (preferably 0.5 or 0.8 in the present invention) to ensure that the algorithm can fully explore different actions in the initial stage of learning. \(\gamma\) represents the decay coefficient, and \(confidence(s t ,t)\) represents the confidence. \(f d (s t )\) represents the update frequency of the data in the database within a set time unit, that is, the ratio of the number of data updates to the time, reflecting the activity of the data. The higher the data update frequency, the more exploration the algorithm needs to adapt to the new data. \(f n (s t )\) represents the network fluctuation, that is, the stability of the network connection within a set time unit (the ratio of the standard deviation of the network delay to the average value of the network delay within the set time unit), reflecting the stability of the network connection. The greater the network fluctuation, the more exploration the algorithm needs to cope with the impact brought by network instability. \(\alpha d \) and \(\alpha n \) both represent weight parameters.
[0058] Dynamically adjusting the exploration probability is an important part of reinforcement learning, especially in the ASRL (Adaptive State Representation Learning) algorithm for adaptive multi-objective optimization, which helps the ASRL algorithm find a balance between exploring new actions and exploiting known best actions.
[0059] S30, in combination with the dynamic exploration probability, select the action of the currently to-be-synchronized data;
[0060] In this embodiment, the specific process of selecting the action of the currently to-be-synchronized data is as follows:
[0061] Generate a random number between 0 and 1;
[0062] If the random number is less than the difference between 1 and the dynamic exploration probability (i.e., 1 - dynamic exploration probability), then randomly select an action from the action set as the action of the currently to-be-synchronized data; otherwise, select the action with the highest Q value in the current state as the action of the currently to-be-synchronized data.
[0063] For the action with the highest Q value in the current state, the acquisition method is:
[0064] Among them, \(a t \) represents the selected action, and \(Q(s t , a; \theta)\) represents at \(s tThe expected return of taking action a, θ represents the parameters of the action-value network (i.e., the Q-network), random_score(a) is the random score, which is used to increase the randomness of exploration. The random score can help the algorithm avoid getting stuck in local optima and ensure sufficient exploration in the early stage. cur(a) is the curiosity reward score, which is used to encourage the algorithm to explore unknown states and actions and can help the algorithm discover new and valuable strategies. h(a) represents the historical score, which reflects the past performance of the action. The historical score can help the algorithm utilize past successful experiences and avoid repeating mistakes. β1, β2, β3, and β4 are all weights.
[0065] S40, execute the selected action, calculate the immediate reward, and observe the state at the next time step; put the current state, the executed selected action, the immediate reward, the state at the next time step, and the importance of the experience into the experience replay pool; sample from the experience replay pool according to the importance of the experience, and update the parameters of the Q-network and the target Q-network;
[0066] In reinforcement learning, the sparsity of the reward signal is a common problem, especially in complex tasks. When the reward signal is sparse (i.e., in most cases, the actions performed by the agent do not immediately produce any positive or negative rewards. Only at specific critical moments or when reaching a certain target state will the agent receive a reward signal. For example, in a maze navigation task, the agent may need to make many moves to reach the end point, and only when reaching the end point will it receive a positive reward), the agent may need to perform many actions to obtain useful feedback, which greatly increases the difficulty of learning. This is specifically manifested in the following aspects:
[0067] 1) Lack of immediate feedback: In a sparse reward environment, it is difficult for the agent to evaluate the quality of its actions through immediate feedback. Since most actions do not immediately produce rewards, it is difficult for the agent to know which actions are beneficial and which are harmful; this lack of immediate feedback makes it difficult for the agent to quickly learn effective strategies through trial and error;
[0068] 2) Inefficient exploration: To find effective strategies, the agent needs to conduct a large amount of exploration. In a sparse reward environment, the exploration process may be very inefficient because the agent needs to try many different action combinations to accidentally discover the path that can obtain rewards;
[0069] 3) Credit assignment problem: Credit assignment refers to attributing the final reward to a series of previous actions. This problem is particularly prominent in sparse reward environments because the agent needs to trace the distant rewards back to actions taken long ago; for example, in a data synchronization task, the agent may need to perform a series of complex synchronization operations to successfully complete a synchronization. If the final reward is only provided after the synchronization is completed, it is difficult for the agent to determine which specific operations played a key role in the successful synchronization;
[0070] 4) Local optimal solution: In a sparse reward environment, the agent is prone to fall into the local optimal solution. Due to the lack of immediate feedback, the agent may find a relatively good solution early on and continue to linger around this local optimal solution without further exploring a better global solution; this situation may cause the agent's learning to stagnate and fail to find the best synchronization strategy.
[0071] In order to solve the above problems. In this embodiment, on the one hand, intermediate rewards are introduced to guide the agent to learn faster. These intermediate rewards can provide positive feedback when the agent approaches the target state, thereby accelerating the learning process; on the other hand, the learning efficiency is improved by encouraging the agent to explore unknown states, and the prediction error of the agent to the environment is calculated to generate intrinsic rewards, thereby motivating the agent to explore new actions and states. Finally, expert demonstrations are used to guide the agent to learn. Experts can provide some successful synchronization examples to help the agent understand and learn effective synchronization strategies faster. The specific formula is as follows:
[0072] Where R represents the immediate reward, I(t) represents the number of data items identified at time step t, and I max represents the maximum number of recognized data items, T(t) represents the amount of data successfully transmitted at time step t, and T max is the maximum amount of data transmitted, ρ(s t , a t ) represents the prediction error, that is, the agent is in state s t Execute a t The error between the predicted value of the reward and the actual reward, d(s t , a t ) indicates that the agent is in state s t Execute action a t Similarity with expert demonstration, ω 识别 ,ω 传输 All represent weight parameters, δ, k, λ, μ all represent attenuation coefficients, η, ψ1 and ψ2 are both positive constants used to adjust the size of the corresponding reward.
[0073] Then, put the current state, the executed selected action, the immediate reward, the state at the next time step, and the importance of the experience into the experience replay pool as an experience;
[0074] Among them, the importance of the experience is calculated based on the TD error of the experience. The TD error reflects the difference between the current experience and the model prediction. A larger TD error means that the experience has a greater potential contribution to the update of the model. The calculation process is as follows: Calculate the difference between the maximum Q value in the next state and the Q value after taking the action in the current state, and take the sum of this difference and the immediate reward in the current state as the TD error. Take the absolute value of the TD error, add it to a very small positive number (to prevent the importance of the experience from being 0), and then perform normalization processing.
[0075] Sample from the experience replay pool according to the importance of the experience (calculate the sampling probability according to the importance of the experience, and sample according to the sampling probability. For example, for 1000 experiences, P represents the importance of the experience, equals 0.6, then 1000 experiences are sampled according to the sampling probability P1), and update the parameters of the Q network and the target Q network, including calculating the target Q value, updating the Q network and the target Q network. Specifically:
[0076] For each sampled sample, calculate the target Q value = discount factor multiplied by the maximum Q value of all actions in the next state predicted by the target Q network + the sum of the immediate rewards corresponding to the samples in the experience replay pool; For example, sample a batch of experiences (s1, a1, 10, s2, p1) from the experience replay pool, which are the current state, the executed selected action, the immediate reward, the state at the next time step, and the importance of the experience respectively, and calculate the target Q value y1: y1 = 10 + 0.9 multiplied by the maximum Q value of all actions in the next state predicted by the target Q network.
[0077] For each sampled sample, calculate the loss function between the predicted value of the Q network and the target Q value (the present invention preferably uses the mean square error loss function. The loss function in traditional Q learning has the following two problems: 1) Since the target Q value is calculated by maximizing the operation, it will lead to an overestimation bias of the Q value; 2) There may be an implicit association between the action values of adjacent states, but the traditional method does not explicitly model it. Therefore, the loss function in this embodiment is, Among them, L represents the loss function, N represents the number of samples, confidence represents the confidence level, respectively represent the predicted value of the Q network and the target Q value, λ L represents the similarity adjustment parameter, sim(s i , s j ) represents the state similarity, Represent the state s i The temporal neighborhood of (i.e., adjacent states), calculate the gradient of the loss function with respect to the Q-network parameters through backpropagation, and update the Q-network parameters using gradient descent method.
[0078] Use the soft update method to gradually update the parameters of the target network based on the parameters θ of the updated Q-network: θ′←τθ+(1-τ)θ′, where τ is the smoothing coefficient.
[0079] In S50, loop through S10 - S40 until the data synchronization reaches the expected performance metrics, then synchronize the data to be synchronized through the trained Q-network.
[0080] In this embodiment, after the Q-network and the target Q-network are trained, load the trained Q-network, obtain the state information of the current data to be synchronized, calculate the dynamic exploration probability, and combine the dynamic exploration probability to select the action of the current data to be synchronized. Additionally, the next state and reward that can be observed can be used to evaluate the effect of the action. If it is found that the environment changes significantly or the synchronization effect is not good, these situations can be recorded for further optimization and adjustment in the future.
[0081] In summary, the present invention can find the best balance between exploration and exploitation by dynamically adjusting the exploration probability, reduce unnecessary resource consumption, and improve the overall performance of the system; and dynamically adjust the weight of the objective function according to the changes in the current environment to ensure that the optimal synchronization strategy can be found in different scenarios, which is applicable to the synchronization tasks of multiple heterogeneous databases and has wide applicability and scalability.
[0082] A heterogeneous database data synchronization system based on reinforcement learning according to the second embodiment of the present invention, the system includes:
[0083] A state acquisition module configured to acquire the state information of the current data to be synchronized as the current state;
[0084] A probability calculation module configured to calculate the arithmetic mean of the difference between the action-value function value and the average action-value function value in the current state based on the current state and the time step as the confidence; combine the confidence and the current state to calculate the dynamic exploration probability: ∈(s t , t)=∈0×exp(-γ×(confidence(s t , t)+α d ×f d (s t )+α n ×f n (s t ));
[0085] Among them, ∈(st, t) represents the dynamic exploration probability, st represents the current state, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the decay coefficient, confidence(st, t) represents the confidence level, fd(st) represents the update frequency of the data in the database within a set time unit, fn(st) represents the network fluctuation, that is, the stability of the network connection within a set time unit, and αd and αn both represent weight parameters;
[0086] An action selection module, configured to select an action for the data to be synchronized currently in combination with the dynamic exploration probability;
[0087] A network update module, configured to execute the selected action, calculate the immediate reward and observe the state of the next time step; put the current state, the executed selected action, the immediate reward, the state of the next time step, and the importance of the experience into the experience replay pool; collect samples from the experience replay pool according to the importance of the experience, and update the parameters of the Q network and the target Q network;
[0088] A data synchronization module, configured to loop through the state acquisition module - the network update module until the data synchronization reaches the expected performance index, and then synchronize the data to be synchronized through the trained Q network.
[0089] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0090] It should be noted that the data synchronization system for heterogeneous databases based on reinforcement learning provided in the above embodiment is only illustrated by dividing the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiment can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present invention, they are only used to distinguish each module or step, and are not regarded as an improper limitation of the present invention.
[0091] An electronic device according to the third embodiment of the present invention includes at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned data synchronization method for heterogeneous databases based on reinforcement learning.
[0092] A computer-readable storage medium according to a fourth embodiment of the present invention, wherein the computer-readable storage medium stores computer instructions for being executed by a computer to implement the above-mentioned data synchronization method for heterogeneous databases based on reinforcement learning.
[0093] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described electronic devices and readable storage media can refer to the corresponding processes in the foregoing method examples and will not be repeated here.
[0094] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0095] The terms "first", "second", "third", etc. are used to distinguish similar objects and not to describe or represent a specific order or sequence.
[0096] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A data synchronization method for heterogeneous databases based on reinforcement learning, characterized in that: The method includes: S10, obtaining status information of the data to be synchronized as the current status; S20, based on the current state and time step, calculate the arithmetic mean of the difference between the action value function value in the current state and the average action value function value as the confidence; combine the confidence and the current state to calculate the dynamic exploration probability: ∈(s t ,t)=∈0×exp(-γ×(confidence(s t ,t)+α d ×f d (s t )+α n ×f n (s t )); Among them, ∈(s t , t) represents the dynamic exploration probability, s t represents the current state, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the attenuation coefficient, and confidence(s t ,t) represents the confidence level, f d (s t ) indicates the update frequency of the data in the database within the set time unit, f n (s t ) represents network fluctuation, that is, the stability of network connection within a set time unit, α d , α n Both represent weight parameters; S30, selecting an action for the data to be synchronized currently in combination with the dynamic exploration probability; S40, executing the selected action, calculating the immediate reward and observing the state of the next time step; putting the current state, the execution of the selected action, the immediate reward, the state of the next time step and the importance of the experience as experience into an experience replay pool; collecting samples from the experience replay pool according to the importance of the experience, and updating the parameters of the Q network and the target Q network; S50, executing S10-S40 in a loop until the data synchronization reaches the expected performance index, and then synchronizing the data to be synchronized through the trained Q network.
2. The data synchronization method for heterogeneous databases based on reinforcement learning according to claim 1 is characterized in that: Combined with the dynamic exploration probability, the action of the data to be synchronized is selected as follows: Generate a random number between 0 and 1; If the random number is less than the difference between 1 and the dynamic exploration probability, an action is randomly selected from the action set as the action of the current data to be synchronized; otherwise, the action with the highest Q value in the current state is selected as the action of the current data to be synchronized.
3. The data synchronization method for heterogeneous databases based on reinforcement learning according to claim 2 is characterized in that: The method for obtaining the action with the highest Q value in the current state is as follows: Among them, a t represents the selected action, Q(s t ,a;θ) represents the t , θ represents the parameters of the action value network, random_score(a) is the random score, cur(a) is the curiosity reward score, h(a) is the historical score, and β1, β2, β3, and β4 are all weights.
4. The data synchronization method for heterogeneous databases based on reinforcement learning according to claim 3 is characterized in that: The instant reward is calculated as follows: Where R represents the immediate reward, I(t) represents the number of data items identified at time step t, and I max represents the maximum number of recognized data items, T(t) represents the amount of data successfully transmitted at time step t, and T max is the maximum amount of data transmitted, ρ(s t ,a t ) represents the prediction error, that is, the agent is in state s t Execute a t The error between the predicted value of the reward and the actual reward, d(s t , a t ) indicates that the agent is in state s t Execute action a t Similarity with expert demonstration, ω 识、 ω 传输 All represent weight parameters, δ, k, λ, μ all represent attenuation coefficients, η, ψ1 and ψ2 are both positive constants used to adjust the size of the corresponding reward.
5. The data synchronization method for heterogeneous databases based on reinforcement learning according to claim 4 is characterized in that: The empirical importance is based on empirical TD error calculations.
6. The method for data synchronization of heterogeneous databases based on reinforcement learning according to claim 5, characterized in that: The method of updating the parameters of the action-value network includes a gradient descent method.
7. A data synchronization system for heterogeneous databases based on reinforcement learning, characterized in that: The system comprises: A status acquisition module is configured to acquire status information of the data to be synchronized as the current status; The probability calculation module is configured to calculate the arithmetic mean of the difference between the action value function value in the current state and the average action value function value based on the current state and the time step as the confidence; and calculate the dynamic exploration probability in combination with the confidence and the current state: ∈(s t ,t)=∈0×exp(-γ×(confidence(s t ,t)+α d ×f d (s t )+α n ×f n (s t )); Among them, ∈(s t , t) represents the dynamic exploration probability, s t represents the current state, t represents the time step, ∈0 represents the initial dynamic exploration probability, γ represents the attenuation coefficient, and confidence(s t , t) represents the confidence, f d (st) indicates the update frequency of the data in the database within the set time unit, f n (s t ) represents network fluctuation, that is, the stability of network connection within a set time unit, α d , α n Both represent weight parameters; An action selection module, configured to select an action for the data to be synchronized currently in combination with the dynamic exploration probability; A network update module is configured to execute a selected action, calculate an immediate reward, and observe a state at a next time step; put the current state, the selected action, the immediate reward, the state at the next time step, and the importance of the experience as experience into an experience replay pool; collect samples from the experience replay pool according to the importance of the experience, and update the parameters of the Q network and the target Q network; The data synchronization module is configured to execute the state acquisition module-the network update module in a loop until the data synchronization reaches the expected performance index, and then synchronize the data to be synchronized through the trained Q network.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor, and a memory communicatively coupled to at least one of the processors; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the data synchronization method of heterogeneous databases based on reinforcement learning as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by a computer to implement the data synchronization method for heterogeneous databases based on reinforcement learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Public opinion evolution analysis method based on multi-agent reinforcement learning
CN116484949A
Heterogeneous database data synchronization method and system
CN117931953A
Multi-agent exploration path planning system based on deep learning
CN119270866A
Machine learning based application of changes in a target database system
US20220284035A1