Data synchronization method and device for distributed computer storage system
By collecting and evaluating the data range, incremental and load characteristics of data synchronization transactions in a distributed computer storage system, formulating node allocation strategies to achieve balanced allocation of data synchronization nodes, solving the delay problem caused by uneven load in traditional systems and improving system performance.
Patent Information
- Application Number
- CN202510107552.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In traditional distributed storage systems, uneven load allocation between storage nodes leads to an increase in delay in data synchronization. Storage nodes may affect synchronization efficiency due to resource overload, thereby slowing down the data synchronization speed of the entire system.
By collecting all data synchronization transactions in the database in the computer storage system, determining their data range and data increments, evaluating the data concentration degree, obtaining the load characteristics and data dependencies of the synchronized storage nodes, formulating node allocation strategies, and dependent quantization of transaction execution conditions based on the policy and data concentration degree to achieve balanced allocation of data synchronization nodes.
The balanced allocation of synchronous storage nodes in the database is realized, which reduces the data delay of the computer storage system in data synchronization, improves system performance, and avoids the situation of node overload or idleness.
Smart Images

Figure CN120086285A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of blockchain distributed storage. More specifically, this application relates to a data synchronization method and device for a distributed computer storage system. Background Art
[0002] Blockchain distributed storage enhances data security and reliability by using the distributed ledger technology of the blockchain in a computer storage system to store data on multiple storage nodes instead of a single data center. The data copies on each node are guaranteed to be consistent by the consensus mechanism of the blockchain, and any tampering behavior on any node will be quickly detected by other nodes, ensuring data integrity and transparency.
[0003] In traditional distributed storage systems, due to differences in hardware configurations such as the storage capacity, computing resources, and network bandwidth of storage nodes, the load distribution among storage nodes is often uneven. Storage nodes may become bottlenecks due to heavy storage tasks, large computing processing loads, or high network transmission pressures, resulting in increased delays during the data synchronization process. For example, some storage nodes may encounter resource overload due to data redundancy, excessive replicas, or centralized processing of specific synchronization tasks, which affects their synchronization efficiency and thus slows down the data synchronization speed of the entire system. Therefore, how to achieve balanced allocation of synchronous storage nodes in the database to reduce data latency in the data synchronization of a computer storage system is a difficult problem faced by the industry. Summary of the Invention
[0004] This application provides a data synchronization method and device for a distributed computer storage system, which can achieve balanced allocation of synchronous storage nodes in the database, thereby reducing data latency in the data synchronization of a computer storage system.
[0005] In a first aspect, this application provides a data synchronization method for a distributed computer storage system, including: When the computer storage system receives a storage request for data synchronization, collect all data synchronization transactions in the database of the computer storage system, and then determine the data range and data increment of each data synchronization transaction; Determine the data concentration of the data to be executed in each data synchronization transaction in the database based on the data range and data increment of each data synchronization transaction; Obtain all synchronous storage nodes in the database of the computer storage system, and determine the node allocation strategy for the database during data synchronization in the computer storage system based on the load characteristics of each synchronous storage node in the database and the data dependency relationship between each data synchronization transaction; Quantify the execution conditions of transactions in the database based on the node allocation strategy and the data concentration of the data executed in each data synchronization transaction to obtain the execution dependencies of the transactions in the database; Based on the execution dependencies, evenly allocate the synchronization storage nodes of each data synchronization transaction in the database.
[0006] In some embodiments, determining the data range and data increment of each data synchronization transaction specifically includes: For each data synchronization transaction, obtain all the database tables operated in the data synchronization transaction; Determine the data range of the data synchronization transaction through all the database tables; Perform incremental verification on the original data in each database table and the data operations in the data synchronization transaction to obtain the data increment of the data synchronization transaction; Furthermore, obtain the data range and data increment of each data synchronization transaction.
[0007] In some embodiments, determining the data concentration of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction specifically includes: For each data synchronization transaction in the database, determine the distribution characteristics of each database table operated in the data synchronization transaction within the data range; Determine the data concentration of the data executed in the data synchronization transaction through all the distribution characteristics and the data increment of the data synchronization transaction, and further obtain the data concentration of the data executed in each data synchronization transaction in the database.
[0008] In some embodiments, determining the node allocation strategy for the database during data synchronization in the computer storage system according to the load characteristics of each synchronization storage node in the database and the data dependency relationship between each data synchronization transaction specifically includes: Determine the dependency priority of each data synchronization transaction in the database through the data dependency relationship between each data synchronization transaction; Determine the allocation priority of each synchronization storage node according to the load characteristics of each synchronization storage node in the database; Generate a node allocation strategy for the database during data synchronization in the computer storage system through each dependency priority and each allocation priority.
[0009] In some embodiments, quantifying the execution conditions of transactions in the database based on the node allocation strategy and the data concentration of the data executed in each data synchronization transaction to obtain the execution dependencies of the transactions in the database specifically includes: For each data synchronization transaction, determine the operation priorities of each database table operated in the data synchronization transaction according to the data concentration degree of the data executed in the data synchronization transaction; Determine the execution priority of the data synchronization transaction through all the operation priorities and the node allocation strategy, and then obtain the execution priorities of each data synchronization transaction; Determine the execution dependency of the transactions in the database according to all the execution priorities.
[0010] In some embodiments, the balanced allocation of the synchronization storage nodes of each data synchronization transaction in the database based on the execution dependency specifically includes: Perform node allocation on all data synchronization transactions based on the data ranges of each data synchronization transaction to obtain the synchronization storage nodes allocated to each data synchronization transaction; Use the execution dependency to perform balanced sorting on all data synchronization transactions in each synchronization storage node to complete the balanced allocation of the synchronization storage nodes in the database.
[0011] In some embodiments, the synchronization storage node is a data node for synchronous storage in a distributed database.
[0012] In a second aspect, the present application provides a data synchronization device for a distributed computer storage system, including: An acquisition module, configured to collect all data synchronization transactions of the database in the computer storage system when the computer storage system receives a storage request for data synchronization, and then determine the data range and data increment of each data synchronization transaction; A processing module, configured to determine the data concentration degree of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction; The processing module is further configured to obtain all synchronization storage nodes of the database in the computer storage system, and determine the node allocation strategy for data synchronization of the database in the computer storage system according to the load characteristics of each synchronization storage node in the database and the data dependency relationship between each data synchronization transaction; The processing module is further configured to perform dependency quantification on the execution conditions of the transactions in the database based on the node allocation strategy and the data concentration degree of the data executed in each data synchronization transaction to obtain the execution dependency of the transactions in the database; An execution module, configured to perform balanced allocation of the synchronization storage nodes of each data synchronization transaction in the database based on the execution dependency.
[0013] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the data synchronization method of the above-mentioned distributed computer storage system.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions or codes are stored. When the instructions or codes run on a computer, the computer is enabled to execute the data synchronization method of the above-mentioned distributed computer storage system.
[0015] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects: In a data synchronization method and device of a distributed computer storage system provided by the present application, when the computer storage system receives a storage request for data synchronization, all data synchronization transactions of the database in the computer storage system are collected, and then the data range and data increment of each data synchronization transaction are determined; the data concentration degree of the data to be executed in each data synchronization transaction in the database is determined through the data range and data increment of each data synchronization transaction; all synchronous storage nodes of the database in the computer storage system are obtained, and a node allocation strategy for the database during data synchronization in the computer storage system is determined according to the load characteristics of each synchronous storage node in the database and the data dependency relationship between each data synchronization transaction; the execution conditions of the transactions in the database are quantitatively dependent according to the node allocation strategy and the data concentration degree of the data to be executed in each data synchronization transaction, and the execution dependency of the transactions in the database is obtained; based on the execution dependency, the synchronous storage nodes of each data synchronization transaction in the database are evenly allocated.
[0016] It can be seen that in this application, the synchronization storage nodes of each data synchronization transaction in the database are evenly allocated based on execution dependency. First, by determining the data concentration, the concentration characteristics of the data range can be obtained. By evaluating the concentration characteristics of the data, tasks with low dependency and high data concentration can be preferentially processed during task scheduling, which helps reduce the latency of data transmission. Thus, the resource requirements of each synchronization transaction can be accurately estimated, unnecessary data synchronization tasks and cross-node data transmissions can be reduced, thereby reducing the overall synchronization latency of the computer storage system and improving the performance of the database in the computer storage system. Then, through a reasonable node allocation strategy, not only can the smooth progress of task execution be ensured, but also the maximized utilization of the resources of the computer storage system can be achieved, avoiding situations of node overload or idleness. Finally, the node allocation strategy can effectively reduce the latency in data synchronization, improve the efficiency and response speed of the computer storage system. By reasonably sorting the execution dependency, the queuing time and waiting time of tasks can be reduced, enabling high-priority synchronization transactions to start as early as possible, thereby reducing latency. At the same time, the priority sorting can also reasonably allocate node resources according to the execution order of transactions, avoiding delays in synchronization tasks in the computer storage system due to insufficient resources. The task scheduling based on execution dependency can dynamically adapt to the load changes of the computer storage system, ensure the balanced execution of synchronization tasks, and avoid delays caused by resource competition or task conflicts. In summary, based on the above solution, the balanced allocation of synchronization storage nodes in the database can be achieved, thereby reducing the data latency in data synchronization of the computer storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is an exemplary flowchart of a data synchronization method for a distributed computer storage system according to some embodiments of the present application; Figure 2 is an architecture diagram of a distributed database according to some embodiments of the present application; Figure 3 is a schematic flowchart of determining execution dependency according to some embodiments of the present application; Figure 4 is a schematic structural diagram of a data synchronization device for a distributed computer storage system according to some embodiments of the present application; Figure 5Schematic diagram of a computer device implementing a data synchronization method for a distributed computer storage system as shown in some embodiments of the present application. Detailed implementation manners
[0019] To better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in conjunction with the specification drawings and specific implementation manners.
[0020] Refer to Figure 1 , which is an exemplary flowchart of a data synchronization method for a distributed computer storage system as shown in some embodiments of the present application. The data synchronization method for the distributed computer storage system mainly includes the following steps: In step 101, when the computer storage system receives a storage request for data synchronization, all data synchronization transactions in the database of the computer storage system are collected, and then the data range and data increment of each data synchronization transaction are determined.
[0021] In some embodiments, refer to Figure 2 As described, this figure is an architecture diagram of a distributed database as shown in some embodiments of the present application. This figure details the architecture of a distributed database, which consists of three core parts: a client layer, a computing service layer, and a storage service layer; the client layer consists of multiple clients, which communicate with the computing service layer through a network; the computing service layer contains multiple query engines, which are responsible for processing requests from clients and interacting with the storage service layer to obtain or store data. The computing service layer also manages the metadata of the data through a metadata service engine, and the metadata service engine includes a metadata manager and a partition manager, which are used to maintain the index and partition information of the data to ensure the fast retrieval and effective management of the data.
[0022] The storage service layer is composed of multiple storage engines, and each storage engine manages several graph partitions, which are the physical storage units of the data. The storage engine organizes the graph partitions through a partition manager to achieve distributed storage of the data. The entire system coordinates the data interaction between the client requests and the storage service through the metadata service engine to ensure the efficient access and processing of the data. In addition, each storage engine is closely connected to the metadata service engine to ensure data consistency and accessibility.
[0023] In some embodiments, the data range and data increment of each data synchronization transaction can be implemented by the following steps: For each data synchronization transaction, obtain all database tables operated in the data synchronization transaction; Determine the data range of the data synchronization transaction through all database tables; Perform incremental verification on the raw data in each database table and the data operations in the data synchronization transaction to obtain the data increment of the data synchronization transaction. Furthermore, obtain the data range and data increment of each data synchronization transaction.
[0024] It should be noted that in this application, the data range refers to the specific data set involved in the data synchronization transaction, usually composed of certain rows and fields in the database table, including all data affected by the synchronization transaction; the data increment refers to the part of the data that has changed in a data synchronization transaction.
[0025] In specific implementation, first, for each data synchronization transaction, obtain all the database tables operated in the data synchronization transaction; second, the set of data operated by the data synchronization transaction in each database table can be used as the data range of the data synchronization transaction; then, for each database table, calculate the data increment of the database table by comparing the database table before and after the data synchronization transaction operation. For update operations, it is necessary to compare the data of each row before and after the operation, find the changed fields and record these changes, and complete the incremental verification of the update operation by comparing based on the primary key and timestamp of the row. For insert operations, record the newly inserted rows, and for delete operations, record the deleted rows. Through the above method, the data increment of each database table can be obtained, and the set of data increments of all database tables can be used as the data increment of the data synchronization transaction; finally, through the above method, the data range and data increment of each data synchronization transaction can be obtained.
[0026] In step 102, determine the data concentration of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction.
[0027] In some embodiments, determining the data concentration of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction can be implemented by the following steps: For each data synchronization transaction in the database, determine the distribution characteristics of each database table operated in the data synchronization transaction within the data range. Determine the data concentration of the data executed in the data synchronization transaction through all the distribution characteristics and the data increment of the data synchronization transaction, and further obtain the data concentration of the data executed in each data synchronization transaction in the database.
[0028] It should be noted that in this application, the data concentration degree represents the concentration degree of the data range in the synchronous transaction; specifically, in implementation, first, for each data synchronous transaction in the database, the amount of data operated in each database table in the data synchronous transaction is counted, so that each amount of data can be used as the distribution characteristic of the corresponding database table within the data range, and this distribution characteristic reflects the distribution state of the data on the database table; then, the product of the maximum value among all distribution characteristics and the data increment of the data synchronous transaction can be used as the data concentration degree of the data executed in the data synchronous transaction. Through the above method, the data concentration degree of the data executed in each data synchronous transaction in the database can be obtained.
[0029] In step 103, all synchronous storage nodes of the database in the computer storage system are obtained, and the node allocation strategy for data synchronization of the database in the computer storage system is determined according to the load characteristics of each synchronous storage node in the database and the data dependency relationship between each data synchronous transaction.
[0030] In some embodiments, obtaining all synchronous storage nodes of the database in the computer storage system can be implemented in the following manner, that is: the synchronous storage nodes can be obtained from the configuration information of the database as all synchronous storage nodes of the database in the computer storage system.
[0031] In some embodiments, determining the node allocation strategy for data synchronization of the database in the computer storage system according to the load characteristics of each synchronous storage node in the database and the data dependency relationship between each data synchronous transaction can be implemented by the following steps: Determine the dependency priority of each data synchronous transaction in the database through the data dependency relationship between each data synchronous transaction; Determine the allocation priority of each synchronous storage node according to the load characteristics of each synchronous storage node in the database; Generate the node allocation strategy for data synchronization of the database in the computer storage system through each dependency priority and each allocation priority.
[0032] In specific implementation, first, the dependency priority of each data synchronization transaction in the database can be determined through the data dependency relationships between various data synchronization transactions, which can be achieved in the following way: it is necessary to identify the dependency relationships between various data synchronization transactions. Among them, the dependency relationships include data dependency, sequential dependency, and cross-node dependency. Data dependency: The execution of a data synchronization transaction depends on the data changes of other transactions. For example, an insert operation may depend on a previous update operation, which means that the insert transaction must be executed after the update transaction. Sequential dependency: When a transaction is executed, a certain order must be ensured. For example, the data in a data table needs to be inserted, updated, or deleted in sequence. Cross-node dependency: When a synchronization transaction operates across multiple storage nodes, it is necessary to clarify the cross-node dependency relationships between transactions, and it is required that the dependent transactions must be executed on the corresponding nodes in a specific order. The topological sorting algorithm can be used to combine the data dependency relationships between various data synchronization transactions to perform priority sorting on the data synchronization transactions, and thus the sorted serial number is used as the dependency priority of each data synchronization transaction in the database. This dependency priority represents the priority order of transaction execution.
[0033] Then, in specific implementation, the allocation priority of each synchronous storage node can be determined according to the load characteristics of each synchronous storage node in the database, which can be achieved in the following way: for each synchronous storage node, evaluate the CPU utilization rate, memory utilization rate, disk I / O load, network bandwidth, and current load of the synchronous storage node as the load characteristics of the synchronous storage node. Among them, CPU utilization rate: whether the computing resources of the storage node are sufficient and the load situation of the computing resources; memory utilization rate: the memory load situation of the synchronous storage node, and the memory bottleneck will affect the execution efficiency of the data synchronization task; disk I / O load: the disk read and write load of the synchronous storage node, and the disk I / O bottleneck may cause synchronization delay; network bandwidth: the network bandwidth of the synchronous storage node, and the network bottleneck will limit the speed of data synchronization; current load: including the number of data synchronization tasks currently being processed by the storage node and the execution time. Thus, based on the load balancing algorithm (such as: round-robin and least connections), data synchronization tasks are allocated for the load characteristics, and the allocation priority of the synchronous storage node is evaluated according to the remaining resources of the synchronous storage node. For example: a synchronous storage node with a lighter load can be preferentially allocated data synchronization tasks and has a higher allocation priority. This allocation priority represents the priority when the synchronous storage node receives data synchronization tasks.
[0034] Finally, in specific implementation, the node allocation strategy for data synchronization of the database in the computer storage system can be implemented by the following method according to each dependency priority and each allocation priority, that is: initialize a node allocation model, use each dependency priority as the weight of the data synchronization task in the node allocation model, use each allocation priority as the weight of the synchronization storage node in the node allocation model, and use the node allocation model to generate the node allocation strategy for data synchronization of the database in the computer storage system. This node allocation strategy is a storage node allocation scheme that ensures the load balance and synchronization efficiency of the system.
[0035] In step 104, according to the node allocation strategy and the data concentration of the data to be executed in each data synchronization transaction, the execution conditions of the transactions in the database are quantified for dependence to obtain the execution dependence of the transactions in the database.
[0036] In some embodiments, according to the node allocation strategy and the data concentration of the data to be executed in each data synchronization transaction, the execution conditions of the transactions in the database are quantified for dependence to obtain the execution dependence of the transactions in the database. Refer to Figure 3 As described above, this figure is a schematic flowchart of determining the execution dependence in some embodiments of the present application. The determination of the execution dependence in this embodiment can be implemented by the following steps: In step 1041, for each data synchronization transaction, determine the operation priority of each database table operated in the data synchronization transaction according to the data concentration of the data to be executed in the data synchronization transaction; In step 1042, determine the execution priority of the data synchronization transaction through all the operation priorities and the node allocation strategy, and then obtain the execution priority of each data synchronization transaction; In step 1043, determine the execution dependence of the transactions in the database according to all the execution priorities.
[0037] It should be noted that in the present application, the execution dependence represents the quantified value of the conditions on which the execution of the transactions in the database depends; the operation priority represents the priority operation order of the database tables in the data synchronization transaction; the execution priority represents the priority execution degree of the data synchronization transaction.
[0038] In specific implementation, first, for each data synchronization transaction, the reciprocal of the data volume of the data concentration degree of the data executed in the data synchronization transaction can be used as the operation priority of each database table operated in the data synchronization transaction; then, the influence weight of each database table on the data synchronization transaction is quantified through a node allocation strategy, so that each influence weight is used as the weight corresponding to the operation priority, and the average value of all operation priorities is calculated as the execution priority of the data synchronization transaction. The execution priorities of each data synchronization transaction can be obtained in the above manner; finally, all the serial numbers after sorting all data synchronization transactions in ascending order according to the corresponding execution priorities can be used as the execution dependency values of each data synchronization transaction in the database, and the set of all execution dependency values can be used as the value range of the execution dependency of the transactions in the database, and the execution dependency of the transactions in the database can be obtained.
[0039] In step 105, based on the execution dependency, the synchronization storage nodes of each data synchronization transaction in the database are evenly allocated.
[0040] In some embodiments, the even allocation of the synchronization storage nodes of each data synchronization transaction in the database based on the execution dependency can be implemented by the following steps: All data synchronization transactions are allocated nodes based on the data range of each data synchronization transaction to obtain the synchronization storage nodes allocated to each data synchronization transaction; The execution dependency is used to evenly sort all data synchronization transactions in each synchronization storage node to complete the even allocation of the synchronization storage nodes in the database.
[0041] In specific implementation, first, all data synchronization transactions are allocated nodes based on the data range of each data synchronization transaction to obtain the synchronization storage nodes allocated to each data synchronization transaction; then, each serial number in the execution dependency is used as the even value for sorting in each synchronization storage node, and all the even values are used to sort all data synchronization transactions in each synchronization storage node to complete the even allocation of the synchronization storage nodes in the database.
[0042] In addition, on the other hand of the present application, in some embodiments, the present application provides a data synchronization device for a distributed computer storage system, referring to Figure 4 , this figure is a schematic structural diagram of a data synchronization device for a distributed computer storage system according to some embodiments of the present application. The data synchronization device for the distributed computer storage system includes: a collection module 201, a processing module 202, and an execution module 203, which are described as follows: The acquisition module 201. In this application, the acquisition module 201 is mainly used to acquire all data synchronization transactions of the database in the computer storage system when the computer storage system receives a storage request for data synchronization, and then determine the data range and data increment of each data synchronization transaction. The processing module 202. In this application, the processing module 202 is used to determine the data concentration degree of the data to be executed in each data synchronization transaction in the database based on the data range and data increment of each data synchronization transaction. It should be noted that the processing module 202 is further used to obtain all synchronous storage nodes of the database in the computer storage system, and determine the node allocation strategy for data synchronization of the database in the computer storage system according to the load characteristics of each synchronous storage node in the database and the data dependency relationship between each data synchronization transaction. In addition, the processing module 202 is further used to perform dependency quantification on the execution conditions of the transactions in the database based on the node allocation strategy and the data concentration degree of the data to be executed in each data synchronization transaction, and obtain the execution dependency of the transactions in the database. The execution module 203. In this application, the execution module 203 is mainly used to evenly allocate the synchronous storage nodes of each data synchronization transaction in the database based on the execution dependency.
[0043] The above has introduced in detail the examples of the data synchronization method and device of the distributed computer storage system provided by the embodiments of this application. It can be understood that, in order to implement the above functions, the corresponding device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0044] In some embodiments, this application further provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned data synchronization method of the distributed computer storage system.
[0045] In some embodiments, refer to Figure 5, the dashed lines in the figure indicate that the unit or module is optional. This figure is a schematic structural diagram of a computer device for implementing the data synchronization method of a distributed computer storage system according to an embodiment of the present application. The data synchronization method of the distributed computer storage system described in the above embodiment can be implemented by Figure 5 the computer device shown. The computer device includes at least one processor 301, a memory 302, and at least one communication unit 305. The computer device can be a terminal device, a server, or a chip.
[0046] The processor 301 can be a general-purpose processor or a dedicated processor. For example, the processor 301 can be a central processing unit (CPU). The CPU can be used to control the computer device, execute software programs, and process the data of software programs. The computer device can also include a communication unit 305 for implementing signal input (reception) and output (transmission).
[0047] For example, the computer device can be a chip, and the communication unit 305 can be the input and / or output circuit of the chip, or the communication unit 305 can be the communication interface of the chip. The chip can be a component of a terminal device, a network device, or other devices.
[0048] Again, for example, the computer device can be a terminal device or a server, and the communication unit 305 can be the transceiver of the terminal device or the server, or the communication unit 305 can be the transceiver circuit of the terminal device or the server.
[0049] The computer device can include one or more memories 302 on which a program 304 is stored. The program 304 can be run by the processor 301 to generate instructions 303, enabling the processor 301 to execute the method described in the above method embodiment according to the instructions 303. Optionally, data (such as a target audit model) can also be stored in the memory 302. Optionally, the processor 301 can also read the data stored in the memory 302. The data can be stored at the same storage address as the program 304, or it can be stored at a different storage address from the program 304.
[0050] The processor 301 and the memory 302 can be set separately or integrated together. For example, they can be integrated on a system on chip (SOC) of a terminal device.
[0051] It should be understood that each step of the above method embodiments can be completed by a logic circuit in the form of hardware or instructions in the form of software in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.
[0052] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0053] For example, in some embodiments, the present application also provides a computer-readable storage medium, in which instructions or code are stored. When the instructions or code run on a computer, the computer is caused to execute the data synchronization method of the above-mentioned distributed computer storage system.
[0054] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0055] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A data synchronization method for a distributed computer storage system, characterized in that: The steps include: When the computer storage system receives a storage request for data synchronization, all data synchronization transactions of the database in the computer storage system are collected, and then the data range and data increment of each data synchronization transaction are determined; Determine the data concentration of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction; Acquire all synchronization storage nodes of the database in the computer storage system, and determine the node allocation strategy of the database when performing data synchronization in the computer storage system according to the load characteristics of each synchronization storage node in the database and the data dependency relationship between each data synchronization transaction; Dependency quantification is performed on the execution conditions of the transactions in the database according to the node allocation strategy and the data concentration of the execution data in each data synchronization transaction to obtain the execution dependency of the transactions in the database; Based on the execution dependency, each data synchronization transaction is evenly distributed among the synchronization storage nodes in the database.
2. The method according to claim 1, characterized in that Determining the data scope and data increment of each data synchronization transaction specifically includes: For each data synchronization transaction, obtain all database tables operated in the data synchronization transaction; Determine the data scope of the data synchronization transaction through all database tables; Perform incremental verification on the original data in each database table and the data operations in the data synchronization transaction to obtain the data increment of the data synchronization transaction; Then the data range and data increment of each data synchronization transaction are obtained.
3. The method according to claim 1, characterized in that Determining the data concentration of the data executed in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction specifically includes: For each data synchronization transaction in the database, determine the distribution characteristics of each database table operated in the data synchronization transaction within the data range; The data concentration of the executed data in the data synchronization transaction is determined through all the distribution characteristics and the data increment of the data synchronization transaction, and then the data concentration of the executed data in each data synchronization transaction in the database is obtained.
4. The method according to claim 1, characterized in that The node allocation strategy for the database when performing data synchronization in the computer storage system is determined according to the load characteristics of each synchronization storage node in the database and the data dependency relationship between each data synchronization transaction, specifically including: Determine the dependency priority of each data synchronization transaction in the database through the data dependency relationship between each data synchronization transaction; Determine the allocation priority of each synchronous storage node according to the load characteristics of each synchronous storage node in the database; A node allocation strategy is generated by various dependency priorities and various allocation priorities when the database performs data synchronization in the computer storage system.
5. The method according to claim 1, characterized in that According to the node allocation strategy and the data concentration of the execution data in each data synchronization transaction, the execution conditions of the transactions in the database are quantified, and the execution dependencies of the transactions in the database are obtained, which specifically include: For each data synchronization transaction, the operation priority of each database table operated in the data synchronization transaction is determined according to the data concentration of the data executed in the data synchronization transaction; Determine the execution priority of the data synchronization transaction through all operation priorities and the node allocation strategy, and then obtain the execution priority of each data synchronization transaction; Determine the execution dependencies of transactions in the database based on all execution priorities.
6. The method according to claim 1, characterized in that Balancing the synchronization storage nodes in the database for each data synchronization transaction based on the execution dependency specifically includes: All data synchronization transactions are assigned nodes based on the data range of each data synchronization transaction to obtain the synchronization storage node assigned to each data synchronization transaction; All data synchronization transactions in each synchronization storage node are evenly sorted using the execution dependency, thereby completing the balanced allocation of synchronization storage nodes in the database.
7. The method according to claim 1, characterized in that The synchronous storage node is a data node that performs synchronous storage in a distributed database.
8. A data synchronization device for a distributed computer storage system, characterized in that: include: A collection module, used to collect all data synchronization transactions of the database in the computer storage system when the computer storage system receives a storage request for data synchronization, and then determine the data range and data increment of each data synchronization transaction; A processing module, used to determine the data concentration of the executed data in each data synchronization transaction in the database through the data range and data increment of each data synchronization transaction; The processing module is also used to obtain all synchronization storage nodes of the database in the computer storage system, and determine the node allocation strategy of the database when performing data synchronization in the computer storage system according to the load characteristics of each synchronization storage node in the database and the data dependency relationship between each data synchronization transaction; The processing module is further used to perform dependency quantification on the execution conditions of the transactions in the database according to the node allocation strategy and the data concentration of the execution data in each data synchronization transaction, so as to obtain the execution dependency of the transactions in the database; The execution module is used to evenly distribute the synchronization storage nodes in the database for each data synchronization transaction based on the execution dependency.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the data synchronization method of a distributed computer storage system according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions or codes, and when the instructions or codes are executed on a computer, the computer implements the data synchronization method of a distributed computer storage system as described in any one of claims 1 to 7.
Citation Information
Cited By
Task execution method and device, computer equipment and storage medium
CN121233667A
Cloud-oriented computer data synchronization method and system
CN122395219A
A cloud-oriented computer data synchronization method and system
CN122395219B