Distributed database synchronization processing method and system, electronic equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]尽管上述系统具备一定的自动化和智能化水平,但在数据同步过程中存在的信息采集不足、性能分析不全面、评估体系不完善等问题,在一定程度上影响了系统的运行效率和同步一致性
[0016]本发明实施例的分布式数据库同步处理方法、系统、电子设备和介质,通过将目标数据库划分为若干子数据区域,并采集各区域对应的同步运行数据,能够实现对分布式环境中数据同步过程的精细化监控。在此基础上,分别计算同步性能动态系数、节点资源性能系数和冲突处理性能系数,并进一步构建综合同步合理度指数,使得同步行为的评估具有全面性和量化基础。通过将该合理度指数与预设标准进行比较,并据此动态调节同步行为,有助于实现对异常同步状态的及时响应与优化控制,从而有效提升数据同步的准确性、稳定性和系统资源的利用效率,增强分布式数据库在多节点复杂场景下的自适应同步能力与一致性保障水平。
Smart Images

Figure CN121166818B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a distributed database synchronization processing method, system, electronic device, and medium. Background Technology
[0002] With the continuous advancement of water resource management concepts, groundwater recharge technology has gradually become an important means to achieve sustainable water resource utilization. Among them, the dynamic simulation and detection method for groundwater recharge based on big data and distributed processing capabilities is the main research method.
[0003] In related technologies, groundwater recharge dynamic simulation and detection systems typically consist of the following functional modules: a data collection unit, a simulation construction unit, a dynamic simulation unit, and an evaluation and early warning unit. By deploying a multi-node data synchronization system, real-time acquisition, simulation, evaluation, and early warning of groundwater recharge-related parameters can be achieved.
[0004] Although the aforementioned systems possess a certain level of automation and intelligence, issues such as insufficient information collection, incomplete performance analysis, and an imperfect evaluation system during data synchronization negatively impact their operational efficiency and synchronization consistency. Furthermore, the systems struggle to accurately grasp the overall state of the distributed data synchronization process and exhibit weak inter-node correlation processing capabilities. These problems are particularly pronounced in complex or large-scale application scenarios, potentially affecting system stability and data consistency. Summary of the Invention
[0005] This invention aims to at least partially solve one of the technical problems in related technologies. Therefore, the objective of this invention is to propose a distributed database synchronization processing method, system, electronic device, and medium for unified control of performance perception and dynamic adjustment of the data synchronization process.
[0006] To achieve the above objectives, a first aspect of the present invention provides a distributed database synchronization processing method, comprising: The target database is divided into several sub-data regions, and data is collected from the sub-data regions to obtain the synchronous operation data corresponding to the sub-data regions. Based on the synchronized operation data, the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient are calculated respectively, and a comprehensive synchronization rationality index is constructed based on the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient. The comprehensive synchronization rationality index is compared with the preset comprehensive synchronization rationality index, and the data synchronization behavior is dynamically adjusted according to the comparison result.
[0007] In addition, the distributed database synchronization processing method of the above embodiments of the present invention may also have the following additional technical features: According to one embodiment of the present invention, the synchronization operation data includes synchronization performance data, which characterizes the data transmission and execution status during the data synchronization task in the sub-data area, and is used for the calculation of the synchronization performance dynamic coefficient. The process of obtaining the synchronous operation data corresponding to the sub-data region includes: The amount of data synchronized per second, the data reception time, the data delay time, and the success rate of synchronization execution are obtained during the process of synchronizing the sub-data areas.
[0008] According to one embodiment of the present invention, the synchronization operation data includes node resource performance data, which characterizes the resource usage of the node during the execution of the data synchronization task of the sub-data region, and is used for the calculation of the synchronization performance dynamic coefficient. The process of obtaining the synchronous operation data corresponding to the sub-data region includes: Obtain the node's processing speed for synchronization tasks, the amount of data used in node memory, the average data load of node I / O, and the node's network bandwidth utilization.
[0009] According to one embodiment of the present invention, the synchronization operation data includes conflict handling data, which characterizes the data consistency during the data synchronization task in the sub-data area, and is used for the calculation of the conflict handling performance coefficient. The process of obtaining the synchronous operation data corresponding to the sub-data region includes: Obtain the number of data conflicts, the time of data conflicts, and the number of successfully resolved conflicts during the synchronization of the sub-data areas.
[0010] According to one embodiment of the present invention, the construction of a comprehensive synchronization rationality index based on synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient includes: The comprehensive synchronization rationality index is determined based on the synchronization performance dynamic coefficient, node resource performance coefficient, conflict handling performance coefficient, and the set weight matrix and influencing factors; the weight matrix consists of the weight of each performance coefficient and the correlation between different performance coefficients.
[0011] According to an embodiment of the present invention, the step of dynamically adjusting the data synchronization behavior based on the comparison result includes: If the comparison result determines that the comprehensive synchronization rationality index is greater than the preset comprehensive synchronization rationality index, an adjustment signal is issued to trigger the adjustment operation on the data synchronization behavior.
[0012] According to one embodiment of the present invention, the method further includes: During the data synchronization process, the sub-data regions are distributed to the corresponding data synchronization endpoints; wherein, the data synchronization endpoint is a combination of a node that issues a synchronization request and the data stored by the node.
[0013] To achieve the above objectives, a second aspect of the present invention provides a distributed database synchronization processing system, comprising: The database partitioning module is used to divide the data in the target database into several sub-data regions; The dataset acquisition module is used to collect data from the sub-data region and obtain the synchronous running data corresponding to the sub-data region; The data synchronization analysis module is used to calculate the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient based on the synchronized operation data. The comprehensive analysis module is used to construct a comprehensive synchronization rationality index based on the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient. The data synchronization adjustment module is used to compare the comprehensive synchronization rationality index with the preset comprehensive synchronization rationality index, and to dynamically adjust the data synchronization behavior based on the comparison result.
[0014] To achieve the above objectives, a third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described distributed database synchronization processing method.
[0015] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the steps of the above-described distributed database synchronization processing method.
[0016] The distributed database synchronization processing method, system, electronic device, and medium of this invention, by dividing the target database into several sub-data regions and collecting synchronization operation data corresponding to each region, can achieve fine-grained monitoring of the data synchronization process in a distributed environment. Based on this, dynamic coefficients for synchronization performance, node resource performance, and conflict handling performance are calculated respectively, and a comprehensive synchronization rationality index is further constructed, giving the evaluation of synchronization behavior a comprehensive and quantitative basis. By comparing this rationality index with preset standards and dynamically adjusting synchronization behavior accordingly, it helps to achieve timely response and optimized control to abnormal synchronization states, thereby effectively improving the accuracy, stability, and system resource utilization efficiency of data synchronization, and enhancing the adaptive synchronization capability and consistency guarantee level of the distributed database in complex multi-node scenarios. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a distributed database synchronization method in one embodiment; Figure 2 This is a block diagram of a distributed database synchronization processing system in one embodiment. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0019] The implementation details of the technical solutions in the embodiments of this application are described in detail below.
[0020] In one embodiment, such as Figure 1 The diagram illustrates a distributed database synchronization method, which may include the following steps: Step S101: Divide the data in the target database into several sub-data areas, collect data from the sub-data areas, and obtain the synchronous running data corresponding to the sub-data areas.
[0021] The target database is divided into several sub-data regions, serving as the basic unit for distributed data synchronization. In practical applications, the data in the target database is denoted as the target data region. To achieve efficient and controllable data synchronization, this target data region is divided into equal parts. Specifically, based on a preset data partitioning strategy, such as according to the number of data records, data size, or logical partitioning, the target data region is divided into several sub-data regions, sequentially labeled as 1, 2, ..., n, for easy identification and processing later. This equal partitioning ensures that each sub-data region bears a relatively balanced load in the synchronization task, improving overall synchronization efficiency while avoiding overload of a single node or uneven resource allocation.
[0022] After the data is partitioned, synchronization-related information is collected for each sub-data region. The synchronization operation data of each region during the actual synchronization process is recorded. The collected synchronization operation data reflects the overall performance of that sub-data region in synchronization behavior. It should be noted that the data collection process here does not directly involve the calculation of specific indicators, but rather serves to establish an initial record of the operational status data exhibited by each sub-region during the actual synchronization process.
[0023] In one embodiment, the synchronization operation data includes synchronization performance data, which characterizes the data transmission and execution during the data synchronization task in the sub-data area, reflecting the transmission efficiency and execution success stability of the sub-data area in the actual data synchronization process. To obtain accurate synchronization performance evaluation results, key performance indicators are recorded in real time during the synchronization of each sub-data area, including but not limited to the amount of data synchronized per second, synchronization data reception time, synchronization data latency, and synchronization execution success rate.
[0024] The amount of data synchronized per second reflects the data transmission rate. This metric is collected through traffic monitoring tools deployed at both the sending and receiving ends. Specifically, data traffic recording modules are set up at both the source and destination nodes to capture the total amount of data transmitted per unit time in real time (e.g., in bytes or data records), and the corresponding data records are updated every second to accurately obtain the amount of data synchronized per second.
[0025] Synchronization data reception time represents the actual timestamp of the receiving node completing data reception, used to analyze the time characteristics of data arrival; when the target node receives synchronization data, it calls a high-precision time acquisition function (such as a microsecond-level or nanosecond-level timestamp function) to record the reception time at this moment, thus obtaining this indicator.
[0026] Synchronization data latency is the difference between the sending and receiving times, used to evaluate the latency characteristics of a network transmission link. At the start of data synchronization, the source node marks each piece of data with a timestamp indicating its generation or transmission. Upon receiving the data, the target node calculates the difference between its own recorded receiving timestamp and the sending timestamp marked by the source node, thus obtaining the transmission latency of that data.
[0027] The synchronization success rate is used to measure the stability of synchronization tasks. Within each statistical period, the number of successfully completed data synchronization tasks and the total number of synchronization tasks attempted to be executed are recorded, and the synchronization success rate for that period is calculated as the ratio of the two, thus reflecting the overall stability and execution reliability of the synchronization process.
[0028] All of the above performance metrics can be extracted during the synchronization process of sub-data areas by recording and analyzing the actual operation events that occur, forming a corresponding synchronization performance data set.
[0029] Based on this, by acquiring and analyzing the data transmission and execution status of the sub-data region, the actual performance of the sub-data region in the data synchronization task can be evaluated more accurately, thereby providing a reliable basis for the calculation of the dynamic coefficient of synchronization performance and improving the accuracy of performance evaluation and the effectiveness of adjustment decisions during the data synchronization process.
[0030] In one embodiment, the synchronized operation data further includes node resource performance data reflecting the running status of the execution nodes. To comprehensively understand the resource usage of the corresponding nodes in each sub-data region during distributed data synchronization, it is necessary to monitor the running status of the nodes performing the synchronization tasks in real time and obtain several representative resource indicators, including node processing speed for synchronization tasks, node memory usage, node I / O load, and node network bandwidth utilization.
[0031] The node processing synchronization task speed is used to measure the node's ability to complete the synchronization task of a sub-data area within a unit of time. During actual data acquisition, the performance statistics interface can be called to obtain the node's CPU time information. Combined with the amount of synchronized data processed by the node and the actual time consumed, the amount of data processed per unit of time can be calculated, thereby evaluating the node's computational processing efficiency.
[0032] Node memory usage reflects the dynamic memory resources used by a node during synchronization tasks. This metric can be queried through the memory management interface to determine the total memory usage of a node, and can also track the specific memory allocation and release of synchronization tasks. These data can be integrated and summarized to form a complete memory usage statistic.
[0033] Node I / O load average reflects the data read / write load between the node and external databases or disk storage. This metric can be collected through integrated disk I / O monitoring tools, which periodically count the total read / write data volume per unit time and calculate the average I / O throughput level to analyze the current node's read / write efficiency and load balancing.
[0034] Node network bandwidth utilization reflects the actual bandwidth usage of the node's network interface during the execution of synchronization tasks. This data is obtained by deploying monitoring software on the node's network interface to capture network traffic data in real time and comparing it with the node's nominal bandwidth capacity, thus revealing the congestion or surplus status of network resources.
[0035] Based on this, by collecting node resource performance data and calculating the dynamic coefficient of synchronization performance, we can more accurately reflect the real-time resource load and processing capacity of each node in the process of executing the synchronization task of the sub-data area, and improve the accuracy and effectiveness of dynamically adjusting the synchronization behavior.
[0036] In one embodiment, the synchronized data further includes data conflict handling information reflecting the risk of data consistency during the synchronization process. Since different nodes may initiate synchronization or write operations on the same data area in a distributed database, data conflicts are highly likely to occur, thus affecting data consistency. Therefore, it is necessary to monitor and record the occurrence of data conflicts in real time, specifically including the number of data conflicts, the time of data conflicts, and the number of successfully resolved conflicts.
[0037] The number of data conflicts measures the frequency of conflict events during synchronization. To obtain this metric, during the execution of sub-data area synchronization tasks, a deployed data conflict detection mechanism monitors concurrent access behavior between nodes in real time and records each detected data conflict. Each time a conflict occurs, a counter automatically increments, and the total number of data conflicts for that sub-data area during the entire synchronization process is obtained through continuous accumulation.
[0038] Data conflict timestamps reflect the temporal characteristics of conflict events. Upon detecting a data conflict, a high-precision time interface is invoked to obtain the moment of occurrence and record it as a conflict timestamp. The obtained time values are saved in a one-to-one correspondence with conflict events, and can be used for subsequent analysis of conflict-intensive time periods, conflict behavior patterns, etc.
[0039] The number of successfully resolved conflicts is used to evaluate the effectiveness of conflict handling during synchronization. After a conflict occurs, the system executes pre-defined conflict resolution logic. Each time a conflict is successfully resolved and data is restored to a consistent state, the success counter is incremented. The number of successfully resolved conflicts can be further used for subsequent synchronization stability assessments or to evaluate the performance coefficient of conflict handling components.
[0040] Based on this, by collecting conflict handling data and calculating the conflict handling performance coefficient, the consistency maintenance efficiency in the data synchronization process can be quantified, enabling the adjustment strategy to be optimized for sub-data areas with frequent conflicts or low processing efficiency, thereby improving the stability and consistency of global data synchronization.
[0041] Step S102: Based on the synchronization operation data, calculate the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient respectively, and construct a comprehensive synchronization rationality index based on the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient.
[0042] To quantitatively evaluate the synchronization effect of each sub-data region, three performance coefficients need to be calculated based on the acquired synchronization operation data: the synchronization performance dynamic coefficient, the node resource performance coefficient, and the conflict handling performance coefficient. These three performance coefficients describe the state of distributed data synchronization from three dimensions: transmission efficiency, node load capacity, and consistency guarantee.
[0043] During synchronization performance analysis, a constructed synchronization performance analysis model can be used to measure the overall performance of each sub-data region during data synchronization. This model takes synchronization performance data from the synchronized data flow as input and calculates dynamic coefficients of synchronization performance. Synchronization performance dynamic coefficient Based on multiple metrics in the synchronization performance data, such as the amount of data synchronized per second, synchronization data reception time, synchronization data latency, and synchronization execution success rate, a multi-layered function combination is used, with the specific expression as follows:
[0044] In the above formula, Indicates the first The amount of data synchronized per second in each sub-data area Indicates the first Synchronization data reception time value for each sub-data area Indicates the first The success rate of synchronizing information in each sub-data area This represents the total number of sub-data regions. The synchronization performance analysis model is based on data throughput, and combines synchronization timing response and success rate stability to construct a weighted relationship, thereby obtaining an overall estimate of synchronization dynamic performance.
[0045] In terms of node resource performance evaluation, a node resource performance coefficient can be calculated based on node resource performance data in the synchronized operation data through a constructed node resource performance analysis model. The specific expression is as follows:
[0046] In the above formula, Indicates the first The speed of node processing synchronization tasks in each sub-data region Indicates the first The amount of data used in the memory of each node in a sub-data region Indicates the first Average data load per node I / O in each sub-data region Indicates the first The node network bandwidth utilization rate of each sub-data region. The numerator of the node resource performance analysis model is used to characterize the non-linear growth trend of processing power, while the denominator normalizes the node memory load. At the same time, by combining the logarithmic and exponential functions of disk load and network occupancy, the actual stress capacity of node resources under synchronous tasks is presented.
[0047] For performance evaluation of synchronization consistency, a conflict handling performance analysis module is constructed, which combines conflict handling data from the synchronization operation data, namely, three indicators: the number of data conflicts, the time of data conflicts, and the number of successfully handled conflicts, to calculate the conflict handling performance coefficient. The specific expression is:
[0048] In the above formula, Indicates the first The number of data conflicts in each sub-data region Indicates the first Data conflict time in individual data regions Indicates the first The number of successful conflict handling operations in each sub-data region. In the conflict handling performance analysis model, the exponential term expresses the amplification effect of the number of successful operations on the overall conflict handling capability. The number of conflicts and the duration of conflicts in the denominator are converted into risk weights in the form of square roots and powers, respectively, so as to accurately evaluate the effectiveness of the synchronous conflict handling mechanism.
[0049] After obtaining the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient, a comprehensive synchronization rationality index is constructed based on these three performance indicators to comprehensively evaluate the overall rationality and efficiency of the current data synchronization behavior. In calculating the comprehensive synchronization rationality index, importance weights for each type of performance indicator and the correlations between different performance indicators can be introduced to perform a fusion calculation of the three types of performance indicators, thereby obtaining a unified reference value that reflects multi-dimensional synchronization performance. In practical applications, the comprehensive synchronization rationality index can reflect not only the stability and efficiency of synchronization but also, to a certain extent, the utilization of system resources and the reliability of conflict handling.
[0050] In one embodiment, to evaluate the overall performance of multiple sub-data regions during the data synchronization process, a comprehensive synchronization analysis model is established, and the synchronization performance dynamic coefficients are used. Node resource performance coefficient and conflict handling performance coefficient As input parameters, the comprehensive synchronization rationality index is calculated. .
[0051] In constructing the comprehensive synchronization analysis model, a weight matrix is introduced to fully reflect the influence of various performance coefficients and their interrelationships. and impact factor The specific expression for the comprehensive synchronous analysis model is:
[0052] in, , representing the vector of comprehensive coefficient values, Representing vectors The transpose of the weight matrix. It includes the weight information of each performance coefficient, and also quantifies the correlation between different performance coefficients, specifically... In the weight matrix The main diagonal elements are fixed at 1, which are used to characterize the dynamic coefficients of synchronization performance. Node resource performance coefficient and conflict handling performance coefficient The importance of each element is reflected in the degree of coupling between the three elements, such as the correlation between synchronization performance and node resources, and the correlation between node resources and conflict handling capabilities.
[0053] Based on the above comprehensive synchronous analysis model, the three types of performance coefficients are fused through functional combination relationships, and combined with influencing factors. The final comprehensive synchronization rationality index is obtained. This is used to comprehensively evaluate the rationality and stability of current distributed data synchronization behavior.
[0054] Step S103: Compare the comprehensive synchronization rationality index with the preset comprehensive synchronization rationality index, and dynamically adjust the data synchronization behavior based on the comparison results.
[0055] After calculating the comprehensive synchronization rationality index, in order to achieve adaptive optimization of the current data synchronization behavior, the calculated comprehensive synchronization rationality index needs to be compared with the system's preset target index. In practical applications, the preset comprehensive synchronization rationality index can be determined based on historical best operating data, empirical rules, or simulation evaluation results, and is used to characterize the synchronization performance level that should be achieved under ideal conditions.
[0056] During the adjustment process, specific synchronization behavior parameters can be adjusted in a targeted manner based on the actual degree of deviation and the dimensions of impact. For example, if a node resource bottleneck is identified, synchronization tasks can be dynamically reallocated, or some sub-data areas can be migrated to idle nodes; if the synchronization performance index is low, the synchronization batch interval can be adjusted, bandwidth allocation can be enhanced, or the synchronization path can be optimized; if the conflict handling coefficient is not ideal, a higher priority locking mechanism or a delayed write caching mechanism can be introduced in data segments with concurrent write operations to improve consistency assurance capabilities.
[0057] In practical applications, dynamic adjustments to data synchronization behavior can be triggered in real time as needed, and a closed-loop adjustment feedback mechanism can be built by combining historical adjustment records. This allows for continuous optimization of the synchronization strategy during operation, ultimately achieving dynamic and adaptive management of data synchronization behavior.
[0058] In one embodiment, after comparing the comprehensive synchronization rationality index with the preset comprehensive synchronization rationality index, if the comparison result shows that the currently calculated comprehensive synchronization rationality index is significantly greater than the preset comprehensive synchronization rationality index, it indicates that the existing data synchronization behavior has anomalies in terms of comprehensive performance, resource utilization and conflict handling, which may lead to decreased efficiency, uneven resource allocation or increased consistency risk.
[0059] At this point, an adjustment signal can be immediately issued based on the comparison results to trigger subsequent adjustments to the data synchronization behavior. In practical applications, this adjustment signal is generated through internal scheduling logic, and subsequently received by administrators who perform manual or semi-automated data optimization. Administrators can select appropriate optimization strategies, such as adjusting the synchronization frequency, reconfiguring the synchronization path, and optimizing task allocation.
[0060] Furthermore, to facilitate timely understanding of system status and adjustment needs by administrators, relevant data adjustment results are displayed visually, including both text and image displays. Text displays showcase specific adjustment suggestions, index comparison values, and descriptions of synchronization anomalies in each sub-data area; image displays, through charts, heatmaps, or flowcharts, intuitively present synchronization status trends, resource distribution, and performance changes before and after adjustments, helping administrators make rapid optimization decisions.
[0061] In one embodiment, to achieve efficient synchronization of sub-data regions in a distributed architecture, during data synchronization, the sub-data regions previously divided according to preset rules need to be distributed to their corresponding data synchronization endpoints. Here, a data synchronization endpoint is represented as a combination of a synchronization node and the data stored on that node, not merely a single physical node; specifically, it is a group. The dataset In this structure, each group For a key-value pair of a synchronization context, where This indicates the identifier of the node that issued the synchronization request. This represents the data currently stored in the node that is used for synchronization; both are representations of token sequences. This approach allows the abstract data synchronization behavior to be modeled as a processable key-value matching relationship.
[0062] In practice, for each sub-data area, a synchronization scheduling mechanism allocates tasks to multiple available nodes in the system. This allocation considers factors such as node resource status, network load, and historical synchronization performance to determine the most suitable target node for synchronizing that sub-data area. The node's performance is recorded during the task allocation process. The unique identifier, and obtain the local data status it holds. Together, they form a logical synchronization endpoint for identifying and executing synchronization operations.
[0063] The aforementioned distributed database synchronization method divides the target database into multiple sub-data regions and collects refined data on the synchronization status of each sub-data region, enabling more granular and comprehensive monitoring of distributed data synchronization behavior. Furthermore, by constructing dynamic synchronization performance coefficients, node resource performance coefficients, and conflict handling performance coefficients respectively, and integrating them into a comprehensive synchronization rationality index, it is possible not only to accurately characterize the overall state of the current synchronization behavior from multiple dimensions, but also to quantitatively identify potential bottlenecks or anomalies. Based on this, and with the help of a preset rationality threshold, dynamic adjustment of data synchronization behavior can be achieved, effectively improving resource utilization, execution stability, and consistency assurance capabilities of the synchronization process.
[0064] In one embodiment, a distributed database synchronization processing system is provided, referencing Figure 2 As shown, the distributed database synchronization processing system 200 may include: a database partitioning module 201, a dataset acquisition module 202, a data synchronization analysis module 203, a comprehensive analysis module 204, and a data synchronization adjustment module 205.
[0065] Among them, the database partitioning module 201 is used to divide the data of the target database into several sub-data regions; The dataset acquisition module 202 is used to collect data from the sub-data area and obtain the synchronous running data corresponding to the sub-data area; The data synchronization analysis module 203 is used to calculate the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient based on the synchronization operation data. The data synchronization analysis module 204 is used to construct a comprehensive synchronization rationality index based on the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient. The data synchronization adjustment module 205 is used to compare the comprehensive synchronization rationality index with the preset comprehensive synchronization rationality index, and to dynamically adjust the data synchronization behavior according to the comparison result.
[0066] In one embodiment, the synchronized running data includes synchronized performance data, which characterizes the data transmission and execution status during the data synchronization task in the sub-data area, and is used for calculating the dynamic coefficient of synchronized performance; the dataset acquisition module 202 is specifically used to acquire the amount of synchronized data per second, the synchronized data reception time, the synchronized data delay time, and the synchronized execution success rate during the synchronized sub-data area process.
[0067] In one embodiment, the synchronized operation data includes node resource performance data, which characterizes the resource usage of the node during the execution of the data synchronization task in the sub-data area, and is used for the calculation of the synchronization performance dynamic coefficient; the dataset acquisition module 202 is specifically used to acquire the node's processing speed of the synchronization task, the amount of data occupied by the node's memory, the amount of data loaded by the node's I / O, and the node's network bandwidth utilization.
[0068] In one embodiment, the synchronized running data includes conflict handling data, which characterizes the data consistency during the data synchronization task in the sub-data area and is used for calculating the conflict handling performance coefficient; the dataset acquisition module 202 is specifically used to acquire the number of data conflicts that occurred during the synchronization of the sub-data area, the time of the data conflicts, and the number of successfully handled conflicts.
[0069] In one embodiment, the data synchronization analysis module 204 is specifically used to determine a comprehensive synchronization rationality index based on the synchronization performance dynamic coefficient, node resource performance coefficient, conflict handling performance coefficient, and a set weight matrix and influencing factors; the weight matrix includes the weight of each performance coefficient and the correlation between different performance coefficients.
[0070] In one embodiment, the data synchronization adjustment module 205 is specifically used to issue an adjustment signal to trigger an adjustment operation on the data synchronization behavior if the comprehensive synchronization rationality index is determined to be greater than the preset comprehensive synchronization rationality index based on the comparison result.
[0071] In one embodiment, the database partitioning module 201 is further configured to distribute sub-data regions to corresponding data synchronization endpoints during the data synchronization process; wherein, a data synchronization endpoint is a combination of a node that issues a synchronization request and the data stored in the node.
[0072] Specific limitations regarding the distributed database synchronization processing system 200 can be found in the limitations of the distributed database synchronization processing method described above, and will not be repeated here. Each module in the aforementioned distributed database synchronization processing system 200 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0073] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement a distributed database synchronization processing method.
[0074] In one embodiment, a computer storage medium is provided on which a computer program is stored, and when the computer program is executed by a processor, it implements a distributed database synchronization processing method.
[0075] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0076] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0077] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0078] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0079] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A distributed database synchronization processing method, characterized in that, include: The target database is divided into several sub-data regions, and data is collected from the sub-data regions to obtain the synchronous operation data corresponding to the sub-data regions. Based on the synchronized operation data, the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient are calculated respectively. A comprehensive synchronization rationality index is then constructed based on these coefficients, including: determining the comprehensive synchronization rationality index according to the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient, as well as a set weight matrix and influencing factors; the weight matrix comprises the weight of each performance coefficient and the correlation between different performance coefficients. The comprehensive synchronization rationality index is compared with a preset comprehensive synchronization rationality index, and the data synchronization behavior is dynamically adjusted based on the comparison result; wherein, The synchronization operation data includes synchronization performance data, which characterizes the data transmission and execution status during the data synchronization task in the sub-data area, and is used for calculating the synchronization performance dynamic coefficient; obtaining the synchronization operation data corresponding to the sub-data area includes: obtaining the amount of data synchronized per second, the synchronization data reception time, the synchronization data delay time, and the synchronization execution success rate during the synchronization of the sub-data area; The synchronization operation data includes node resource performance data, which characterizes the resource usage of the node during the execution of the data synchronization task in the sub-data region, and is used for the calculation of the synchronization performance dynamic coefficient; obtaining the synchronization operation data corresponding to the sub-data region includes: obtaining the node's processing speed of the synchronization task, the amount of data occupied by the node's memory, the amount of data loaded by the node's I / O, and the node's network bandwidth utilization rate. The synchronization operation data includes conflict handling data, which characterizes the data consistency during the data synchronization task in the sub-data area and is used for calculating the conflict handling performance coefficient. Obtaining the synchronization operation data corresponding to the sub-data area includes obtaining the number of data conflicts, the time of data conflicts, and the number of successfully handled conflicts during the synchronization of the sub-data area.
2. The distributed database synchronization processing method according to claim 1, characterized in that, The dynamic adjustment of data synchronization behavior based on comparison results includes: If the comparison result determines that the comprehensive synchronization rationality index is greater than the preset comprehensive synchronization rationality index, an adjustment signal is issued to trigger the adjustment operation on the data synchronization behavior.
3. The distributed database synchronization processing method according to claim 1, characterized in that, The method further includes: During the data synchronization process, the sub-data regions are distributed to the corresponding data synchronization endpoints; wherein, the data synchronization endpoint is a combination of a node that issues a synchronization request and the data stored by the node.
4. A distributed database synchronization processing system, characterized in that, For implementing the distributed database synchronization processing method according to any one of claims 1 to 3, the distributed database synchronization processing system comprises: The database partitioning module is used to divide the data in the target database into several sub-data regions; The dataset acquisition module is used to collect data from the sub-data region and obtain the synchronous running data corresponding to the sub-data region; The data synchronization analysis module is used to calculate the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient based on the synchronized operation data. The comprehensive analysis module is used to construct a comprehensive synchronization rationality index based on the synchronization performance dynamic coefficient, node resource performance coefficient, and conflict handling performance coefficient. The data synchronization adjustment module is used to compare the comprehensive synchronization rationality index with the preset comprehensive synchronization rationality index, and to dynamically adjust the data synchronization behavior based on the comparison result.
5. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the distributed database synchronization processing method according to any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the distributed database synchronization processing method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Database data real-time synchronization method and device
CN119202076A
System and method for decentralized online data transfer and synchronization
US20130179947A1