Database processing method and device, first node and readable storage medium

When the vector clock of the first node of the distributed database system is updated, whether to prune is determined based on the time difference and clock length, the problem of increased vector clock space occupancy is solved, data transmission and processing efficiency is improved, and system performance is improved.

CN120045618APending Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311602067.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In a distributed database system, the length of the vector clock is positively correlated with the number of nodes, resulting in an increase in space occupancy and reducing data transmission efficiency and database algorithm processing efficiency.

Method used

When the vector clock of the first node is updated, it is decided whether to prune the vector clock based on the time difference between the earliest update time and the current time and/or the clock length. During the pruning process, the value of the earliest updated dimension is removed from the vector clock values ​​of multiple dimensions to obtain the pruned vector clock.

Benefits of technology

On the premise of ensuring the accuracy of vector clock, the space occupation of vector clock is effectively reduced, data transmission efficiency and database algorithm processing efficiency are improved, thereby improving the overall processing efficiency and performance of distributed database systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045618A_ABST
    Figure CN120045618A_ABST
Patent Text Reader

Abstract

The invention discloses a database processing method and device, a first node and a readable storage medium. The processing efficiency and performance of a distributed database system can be improved. Comprising the following steps: under the condition that a first vector clock corresponding to a first node is updated, determining whether to prune the first vector clock based on the time difference between the earliest updating time corresponding to the first vector clock and the current time and / or the clock length of the first vector clock; the first vector clock comprises first vector clock values of at least two dimensions; the earliest update time is the update time corresponding to the earliest update dimension in the at least two dimensions; and when it is determined that the first vector clock is pruned, removing the first vector clock value of the dimension updated earliest from the first vector clock values of the at least two dimensions, determining the pruned first vector clock, and updating the first vector clock by using the pruned first vector clock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of databases, and in particular, to a database processing method, apparatus, first node, and readable storage medium. Background Art

[0002] Currently, the clock technologies in distributed database systems mainly include: physical clock technology, logical clock technology, and hybrid clock technology. All of these clock technologies have the problem of inaccurate time. In order to improve the clock accuracy, vector clock technology has been developed on the basis of logical clock technology. The accuracy of vector clock technology has been improved, but the length of the vector clock is positively correlated with the number of nodes in the distributed system. Therefore, the clock length may continuously increase as the number of nodes increases, resulting in an increase in the space occupancy rate of the vector clock, reducing the data transmission efficiency related to the vector clock and the processing efficiency of database algorithms, and thus reducing the processing efficiency and performance of the distributed database system. Summary of the Invention

[0003] Embodiments of this application are expected to provide a database processing method, apparatus, first node, and readable storage medium that can improve the processing efficiency and performance of a distributed database system.

[0004] The technical solution of this application is implemented as follows:

[0005] In a first aspect, an embodiment of this application provides a database processing method applied to a first node in a distributed system. The method includes:

[0006] When the first vector clock corresponding to the first node is updated, determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock; the first vector clock includes first vector clock values in at least two dimensions; the earliest update time is the update time corresponding to the earliest updated dimension among the at least two dimensions;

[0007] When it is determined to prune the first vector clock, remove the first vector clock value of the earliest updated dimension from the first vector clock values in the at least two dimensions to complete the pruning of the first vector clock.

[0008] In a second aspect, an embodiment of this application provides a database processing apparatus applied to a first node in a distributed system. The apparatus includes:

[0009] A determination module, configured to determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock, when the first vector clock corresponding to the first node is updated; the first vector clock includes first vector clock values in at least two dimensions; the earliest update time is the update time corresponding to the dimension that is updated earliest among the at least two dimensions;

[0010] A pruning module, configured to, when it is determined to prune the first vector clock, remove the first vector clock value of the earliest updated dimension from the first vector clock values in the at least two dimensions, so as to complete the pruning of the first vector clock.

[0011] In a third aspect, an embodiment of the present application provides a first node, including a memory and a processor; wherein,

[0012] The memory is configured to store executable instructions;

[0013] The processor is configured to implement the database processing method provided by the embodiment of the present application when executing the executable instructions stored in the memory.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, storing executable instructions, which are used to cause a processor to implement the database processing method provided by the embodiment of the present application when executed.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, where the computer program or instruction implements the database processing method provided by the embodiment of the present application when executed by a processor.

[0016] The present application provides a database processing method, apparatus, first node, and readable storage medium. When the first vector clock corresponding to the first node is updated, it is determined whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock. Among them, the first vector clock includes first vector clock values in at least two dimensions, and the earliest update time is the update time corresponding to the dimension with the earliest update among at least two dimensions. When it is determined to prune the first vector clock, the first vector clock value of the earliest updated dimension is removed from the first vector clock values in at least two dimensions, the pruned first vector clock is determined, and the first vector clock is updated using the pruned first vector clock. The present application realizes determining whether to prune the vector clock according to at least one of the clock length and the earliest update time of the vector clock, effectively reducing the space occupied by the vector clock while ensuring the accuracy of the vector clock, thereby improving the data transmission efficiency related to the vector clock and the processing efficiency of database algorithms, and further improving the processing efficiency and performance of the distributed database system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an optional flowchart of the database processing method provided by an embodiment of the present application;

[0018] Figure 2 is an optional flowchart of the database processing method provided by an embodiment of the present application;

[0019] Figure 3 is an optional schematic diagram of a pruning area on a two-dimensional coordinate system composed of the clock length and the time difference provided by an embodiment of the present application;

[0020] Figure 4 is a schematic diagram of the organizational structure of queue elements of a small root heap structure queue provided by an embodiment of the present application;

[0021] Figure 5 is a schematic diagram of the process of updating the order of queue elements of a small root heap structure queue provided by an embodiment of the present application Figure 1 ;

[0022] Figure 6 is a schematic diagram of the process of updating the order of queue elements of a small root heap structure queue provided by an embodiment of the present application Figure 2 ;

[0023] Figure 7 is a schematic diagram of the process of updating the update time corresponding to the dimension that has been updated in a small root heap structure queue provided by an embodiment of the present application;

[0024] Figure 8Schematic diagram of the process of adjusting the order and update time of queue elements in the small root heap structure queue provided by the embodiment of the present application;

[0025] Figure 9 Schematic diagram of the comparison of the space occupied by the vector clock before and after pruning provided by the embodiment of the present application;

[0026] Figure 10 An optional flowchart of the database processing method provided by the embodiment of the present application;

[0027] Figure 11 An optional structural diagram of the database processing device provided by the embodiment of the present application;

[0028] Figure 12 An optional structural diagram of a first node provided by the embodiment of the present application. Detailed implementation manners

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0030] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0031] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0033] Currently, the clock technologies in distributed database systems mainly include: physical clock technology, logical clock technology, and hybrid clock technology. For physical clock technology, the current technical solution can be that the node servers in the distributed system are responsible for processing multiple transaction requests. If concurrent conflicts are found among transactions, the physical clock method is used to obtain the start time of each transaction. The calculation and calibration of the physical clock are achieved by synchronizing the local time of the node server with the time of the Global Positioning System (GPS) atomic clock server in its location area, and then corresponding processing is performed on each transaction. Among them, each transaction is assigned a unique timestamp, and the start time and commit time of the transaction are recorded on each node. To reduce the deviation value of the physical clock between different nodes, a GPS atomic clock is used to provide a high-precision time signal, which is used as the global clock to coordinate the time of each node in the distributed database system.

[0034] For logical clock technology, the current technical solution can characterize the time order between events through logical time values. The nodes in the distributed system can include three types - coordinator nodes, catalog nodes, and data nodes. When a transaction starts, the transaction start time is set to the local logical time of the coordinator node; if the difference value between the local logical time of the data node and the coordinator node exceeds a preset threshold, the local logical time of the data node is calibrated, and the coordinator node rolls back the transaction. After calibrating the time of the coordinator node, the transaction is retried; if the difference value between the transaction pre-commit time and the transaction start time is greater than the transaction tolerance error of the transaction, the transaction pre-commit time and the pre-commit message of the transaction are sent to all data nodes participating in the transaction; according to the difference value situation of the timestamps of two different transactions, a data node is selected as the arbitration node to arbitrate the timestamp order of the two different transactions.

[0035] For hybrid clock technology, each transaction is assigned a unique Hybrid Logic Clock (HLC) timestamp and is scheduled and executed in chronological order. By comparing the timestamps of transactions, the system can ensure the sequentiality and consistency of transactions, avoiding conflicts and data inconsistency problems caused by concurrent execution. The key technology to achieve data consistency uses HLC and its algorithm to describe the time order between events. HLC technology combines the physical clock and the logical clock, and simultaneously stores the global physical clock information and logical clock information when the data was last updated. The recorded global physical clock value must satisfy being greater than or equal to the local physical clock value of the current node.

[0036] It can be seen that for the current physical clock technology, due to the clock drift problem of the physical clock between different nodes, the frequencies of the clocks may be different, resulting in inaccurate time and being easily affected by external factors. Therefore, the hardware configuration requirements for devices are relatively high, otherwise it is difficult to ensure the accuracy of time.

[0037] For the current logical clock technology, since the logical clock cannot provide a global clock order and can only provide a partial order relationship between events, in some cases, the logical clock cannot accurately judge the event time sequence. Therefore, the time accuracy of the logical clock is also relatively low.

[0038] For the current hybrid clock technology, the hybrid clock technology combines the characteristics of physical clocks and logical clocks. Therefore, the complexity in design and implementation is higher, the maintenance cost is also greater, and certain system overhead may be introduced. At the same time, due to the same dependence on physical clocks, it is difficult to guarantee the high precision and reliability of time.

[0039] To improve the clock accuracy, vector clock technology has been developed based on the current logical clock technology. Vector clock technology is an extension of logical clock technology. The accuracy of vector clocks is higher than that of the above several clock technologies, but vector clocks have the disadvantage that the clock length may grow infinitely. Since the length of the vector clock is positively correlated with the number of nodes in the distributed system, when the number of nodes in the cluster is large, the length of the vector clock will also be long, resulting in an increase in the space occupancy rate of the vector clock. An overly large space size is not conducive to accelerating the processing of vector clocks in comparison algorithms, nor is it conducive to the communication process of vector clock data between nodes, thus affecting the performance of the entire consistency protocol and further reducing the processing efficiency and performance of the distributed database system.

[0040] The embodiments of the present application provide a database processing method, apparatus, first node, and readable storage medium, which can improve the flash calibration efficiency. See Figure 1 , Figure 1 This is an optional flowchart of the database processing method provided by the embodiments of the present application. As follows:

[0041] S101. When the first vector clock corresponding to the first node is updated, determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock.

[0042] The database processing method of the embodiments of the present application is applied to the database processing of a distributed system. The distributed system includes at least two nodes, and the at least two nodes at least include a first node and a second node. The distributed system in the embodiments of the present application uses vector clock technology to ensure the orderliness of operations and data consistency.

[0043] In the embodiments of the present application, the vector clock corresponding to each node in the distributed system includes at least two vector clock values corresponding to at least two nodes in the distributed system. Exemplarily, the distributed system includes n nodes, where n is an integer greater than 1. The vector clock corresponding to the i-th node among the n nodes includes Vi[1], Vi[2],..., Vi[n]. Among them, Vi[1] is the vector clock value on the first node known to the i-th node; Vi[2] is the vector clock value on the second node known to the i-th node; and so on, Vi[n] is the vector clock value on the n-th node known to the i-th node. In this way, the i-th node can understand the state of the data replicas on each node in the distributed system according to the vector clock maintained by itself.

[0044] In some embodiments, the vector clock value represents the version information of the data replica on the corresponding node. Exemplarily, the vector clock value may include a timestamp of the local time of the corresponding node, or include an ordered number generated by the corresponding node according to a preset rule, etc. The specific selection is made according to the actual situation, and the embodiments of the present application do not make limitations. Exemplarily, the initial value of the vector clock corresponding to each node may be 0. Each time data is updated on this node, among the at least two vector clock values corresponding to the at least two nodes maintained by this node, the vector clock value corresponding to itself increases by a preset step. That is to say, the larger the vector clock value, the newer the data version.

[0045] In the embodiments of the present application, the first node may be any one of at least two nodes in the distributed system. The first vector clock is the vector clock corresponding to the first node. The first vector clock includes the vector clock values of each node in the distributed system known to the first node. In some embodiments, the first vector clock includes first vector clock values in at least two dimensions; the at least two dimensions correspond to at least two nodes in the distributed system.

[0046] In the embodiments of the present application, when data synchronization is performed between the first node and the node to be synchronized in the distributed system, the first vector clock corresponding to the first node is compared with the vector clock corresponding to the node to be synchronized. Here, when at least one of the first clock vector values in the at least one dimension of the first vector clock corresponding to the first node is less than the vector clock value corresponding to the at least one dimension in the node to be synchronized, it indicates that the data version of the first node in the at least one dimension is older than that of the node to be synchronized, and the data file of the first node in the at least one dimension needs to be updated. When the first node determines to update the data file in the at least one dimension, the first vector clock corresponding to the first node is updated. The first node updates the first vector clock value corresponding to the at least one dimension in the first vector clock according to the vector clock value corresponding to the at least one dimension in the node to be synchronized. Moreover, the first node records the time of the last update of the first vector clock value corresponding to each of at least two dimensions as the update time corresponding to each of at least two dimensions.

[0047] In the embodiments of the present application, when the first vector clock corresponding to the first node is updated, the first node determines whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock.

[0048] Regarding the problem of infinite growth of space existing in the current vector clock technology, the embodiments of the present application can remove the stale information that has the least impact on the data version in the vector clock through a pruning algorithm, thereby reducing the space occupied by the vector clock. Since the vector clock is a data structure that identifies data version information, in order to avoid reducing the reliability of the data version due to removing some information of the vector clock, the embodiments of the present application make a trade-off based on the update frequency and / or clock length of the vector clock to reduce the space occupied by the vector clock on the basis of ensuring the reliability of the data version information.

[0049] In some embodiments, when the time difference between the earliest update time corresponding to the first vector clock and the current time is greater than or equal to the preset time difference threshold, it is determined to prune the first vector clock; when the time difference is less than the preset time difference threshold, it is determined not to prune the first vector clock.

[0050] In the embodiments of the present application, the first vector clock includes first vector clock values of at least two dimensions; the vector clock value of each dimension represents the last update time of the node corresponding to that dimension. The earliest update time corresponding to the first vector clock is the update time corresponding to the dimension with the earliest update among at least two dimensions. That is to say, the earliest update time is the update time corresponding to the dimension with the earliest last update time among at least two dimensions.

[0051] In the embodiments of the present application, when the first vector clock is updated, if the time difference between the earliest update time corresponding to the first vector clock and the current time is less than the preset time difference threshold, it indicates that among the first vector clock values of at least two dimensions included in the first vector clock, the maximum time difference from the current time is less than the preset time difference threshold, that is, the first vector clock is in a state of frequent update. For a vector clock with frequent updates, removing the information of some dimensions may cause the data version of the data copy on the first node to be inconsistent with that of other nodes, reducing the credibility of the data copy. Therefore, to ensure data reliability, when the time difference between the earliest update time corresponding to the first vector clock and the current time is greater than or equal to the preset time difference threshold, the first node determines to prune the first vector clock. When the time difference is less than the preset time difference threshold, it is determined not to prune the first vector clock.

[0052] In some embodiments, when the clock length of the first vector clock is greater than or equal to the preset clock length threshold, it is determined to prune the first vector clock; when the clock length is less than the preset clock length threshold, it is determined not to prune the first vector clock.

[0053] In the embodiments of the present application, when the clock length of the first vector clock is greater than or equal to the preset clock length threshold, it indicates that the space occupied by the first vector clock is too large, which may affect the data processing efficiency and data transmission efficiency. The first node determines to prune the first vector clock; when the clock length is less than the preset clock length threshold, it is determined not to prune the first vector clock.

[0054] In some embodiments, the first node can determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and the clock length of the first vector clock.

[0055] In the embodiments of the present application, the first node can combine two factors, namely the time difference between the earliest update time and the current time and the clock length, to determine whether to prune the first vector clock.

[0056] In some embodiments, in the design of the pruning algorithm, the priority of the time difference can be higher than that of the clock length. Exemplarily, when the clock length is greater than or equal to the preset clock length threshold, it is determined whether the time difference is greater than or equal to the preset time difference threshold. When the clock length is greater than or equal to the preset clock length threshold and the time difference is greater than or equal to the preset time difference threshold, it is determined to prune the first vector clock. When the clock length is less than the preset clock length threshold, or the time difference is less than the preset time difference threshold, it is determined not to prune the first vector clock.

[0057] In an embodiment of the present application, when the clock length is greater than or equal to a preset clock length threshold, the first node further determines whether the time difference between the earliest update time corresponding to the first vector clock and the current time is greater than or equal to a preset time difference threshold. When the clock length is greater than or equal to the preset clock length threshold and the time difference is greater than or equal to the preset time difference threshold, it is determined to prune the first vector clock. When the clock length is greater than or equal to the preset clock length threshold and the time difference is less than the preset time difference threshold, it is determined not to prune the first vector clock. When the clock length is less than the preset clock length threshold, it is no longer further determined whether the time difference is greater than or equal to the preset time difference threshold, and it is directly determined not to prune the first vector clock.

[0058] S102. When it is determined to prune the first vector clock, remove the first vector clock value of the earliest updated dimension from the first vector clock values of at least two dimensions to complete the pruning of the first vector clock.

[0059] In an embodiment of the present application, when the first node determines to prune the first vector clock, it removes the first vector clock value of the earliest updated dimension from the first vector clock values of at least two dimensions, thereby completing the pruning of the first vector clock and obtaining the pruned first vector clock. The first node updates the first vector clock with the pruned first vector clock, that is, uses the pruned first vector clock as the first vector clock. In this way, the space occupied by the first vector clock is reduced through pruning.

[0060] Exemplarily, in an actual application scenario, the vector clock pruning process in the database management method provided by the embodiment of the present application can be as Figure 2 shown, including S201 - S206, as follows:

[0061] S201. Update the first vector clock.

[0062] In S201, when the first node receives a data synchronization request sent by a node to be synchronized in the distributed system and determines to update the data file on the first node, the first node correspondingly updates the first vector clock.

[0063] S202. Determine whether the clock length is greater than or equal to a preset clock length threshold.

[0064] In S202, when the first vector clock is updated, the first node determines whether the clock length of the first vector clock is greater than or equal to the preset clock length threshold. If so, execute S204; if not, execute S203.

[0065] S203. Do not prune the first vector clock.

[0066] In S203, if the clock length of the first vector clock is less than the preset clock length threshold, the first node does not prune the first vector clock. Alternatively, when the time difference of the update of dimension j is less than the preset time difference threshold, the first node does not prune the first vector clock.

[0067] S204. Obtain the currently oldest dimension identifier j.

[0068] In S204, if the clock length of the first vector clock is greater than or equal to the preset clock length threshold, the first node uses the dimension with the earliest last update time among at least two dimensions corresponding to the first vector clock as the currently oldest dimension, and obtains the currently oldest dimension identifier j.

[0069] S205. Determine whether the time difference of the update of dimension j is greater than or equal to the preset time difference threshold.

[0070] In S205, the time difference of the update of dimension j is the time difference between the last update time of dimension j and the current time. The first node determines whether the time difference of the update of dimension j is greater than or equal to the preset time difference threshold. If so, execute S206; if not, execute S203.

[0071] S206. Remove the vector clock value corresponding to dimension j.

[0072] In S206, the first node removes the vector clock value corresponding to dimension j, thereby completing one pruning of the first vector clock and reducing the space occupied by the first vector clock.

[0073] It can be seen that in the vector clock pruning process of the database processing method according to the embodiments of the present application, in fact, according to the clock length of the first vector clock and the time difference between the last update time of the first vector clock in the oldest dimension and the current time, a feature point of the first vector clock is formed, and the feature point is merged on the two-dimensional coordinate system composed of the clock length and the time difference. If the feature point falls into the pruning area as shown in Figure 3 shown, the pruning operation will be performed on the first vector clock; otherwise, no pruning operation will be performed. As shown in Figure 3 shown, when the time difference is small, such as when the time difference is less than the preset time difference threshold, even if the clock length of the first vector clock has exceeded the preset clock length threshold, the first node will not perform pruning to reduce the space occupied by the first vector clock to ensure the reliability of the data version information on the first node.

[0074] It can be understood that, when the first vector clock corresponding to the first node is updated in the embodiments of the present application, it is determined whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock. When it is determined to prune the first vector clock, the first vector clock value of the earliest updated dimension is removed from the first vector clock values in at least two dimensions, the pruned first vector clock is determined, and the first vector clock is updated using the pruned first vector clock. The present application realizes determining whether to prune a vector clock based on at least one of the clock length and the earliest update time of the vector clock, effectively reducing the space occupied by the vector clock while ensuring the accuracy of the vector clock, thereby improving the data transmission efficiency related to the vector clock and the processing efficiency of the database algorithm, and further improving the processing efficiency and performance of the distributed database system.

[0075] In some embodiments, before the first node determines whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, or before the first node determines whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and the clock length of the first vector clock, the earliest update time can be obtained from a preset queue. Here, the preset queue is used to record the time of the last update of the first vector clock value corresponding to each dimension in at least two dimensions.

[0076] In the embodiments of the present application, when the first node records the time of the last update of the first vector clock value corresponding to each dimension in at least two dimensions as the update time corresponding to each dimension in at least two dimensions, it will not combine the update time corresponding to each dimension with the data structure of the first vector clock itself, because if implemented in this way, it may instead increase the data volume of the vector clock. The first node can record the time of the last update of the first vector clock value corresponding to each dimension in at least two dimensions through a dedicated preset queue, and when it is necessary to determine whether to prune the first vector clock based on the time difference, obtain the earliest update time corresponding to the first vector clock from the preset queue, and then determine the time difference based on the earliest update time and the current actual time.

[0077] In some embodiments, the preset queue is a small root heap structure queue, and the first node can obtain the earliest update time from the first queue element of the small root heap structure queue.

[0078] In the embodiments of the present application, based on the characteristics of the small root heap data structure, the value of the first queue element in the small root heap structure queue is the smallest. Therefore, the update time recorded by the first queue element of the small root heap queue is the earliest update time corresponding to the first vector clock.

[0079] In some embodiments, when it is determined to prune the first vector clock, the first node removes the dimension corresponding to the current first queue element from the small root heap structure queue, thereby updating the small root heap structure queue to ensure the data structure characteristics of the small root heap. Exemplarily, the first node deletes the current first queue element from the small root heap structure queue, and uses the current second queue element as the first queue element after the update of the small root heap structure queue, and so on, to implement the update of the small root heap structure queue. Among them, the pruning of the first vector clock by the first node and the update of the small root heap structure queue can be executed synchronously, or can be executed in any order, and specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.

[0080] In some embodiments, when the first vector clock is updated, the first node determines the dimension where the update occurs; according to the current time, the update time corresponding to the dimension where the update occurs in the small root heap structure queue is updated, and the order of the queue elements in the small root heap structure queue is updated to ensure the data structure characteristics of the small root heap. Exemplarily, when the first node updates the update time corresponding to the dimension where the update occurs in the small root heap structure queue to the current time, the current queue element corresponding to the dimension where the update occurs is adjusted to the last queue element of the small root heap structure queue, and the queue elements after the current queue element are moved forward in turn to implement the update of the small root heap structure queue. Among them, the update of the first vector clock by the first node and the update of the small root heap structure queue can be executed synchronously, or can be executed in any order, and specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.

[0081] In some embodiments, the small root heap structure queue can be as Figure 4 shown, including the update time T corresponding to dimension A A to the update time T corresponding to dimension H H , and dimensions A to H respectively correspond to the respective queue elements in the small root heap structure queue. In the Figure 4 shown small root heap structure queue, the value of the parent node is less than the value of its child node. Among them, dimension A is the dimension corresponding to the first queue element (heap top element) of the small root heap structure queue. Taking dimension C as an example, Figure 4 the update time corresponding to dimension C in it is T C , if dimension C is updated, the update time corresponding to dimension C after the update is T C ', the process by which the first node updates the update time corresponding to the dimension where the update occurs in the small root heap structure queue according to the current time and updates the order of the small root heap structure queue may include: 1. As Figure 5As shown, swap the order of dimension C and dimension H of the last queue element (heap tail element) of the small root heap structure queue. 2. Update the update time T corresponding to dimension C before the update C with the update time T corresponding to dimension H H for comparison. Since dimension C is the parent node of dimension F and dimension G, so T C is less than T F and T G the minimum value in, that is, T C <min{T F ,T G}. If T H >T C , then perform a sink operation on dimension H: in T F <T G case, continue to compare T H with T F , if T H >T F , then as Figure 6 shown, swap dimension H and dimension F; otherwise do not swap. If T H <T C , then perform a float operation on dimension H: compare T H with the update time T corresponding to the current parent node dimension A of dimension H A for comparison, if T H <T A , then swap T H with T A , otherwise do not swap, thereby updating the order of the queue elements. If T H =T C , then keep the order after swapping dimension C and dimension H, and no further swapping is required. 3. Update the update time corresponding to the current heap tail element, that is, dimension C, to T C ', as Figure 7 shown.

[0082] In some embodiments, the process of the first node adjusting the order of the queue elements of the small root heap structure queue and the update time can be as Figure 8 shown, including:

[0083] S301. Determine that the value of a certain dimension D of the vector clock i is updated at time t new .

[0084] S302. Determine whether D i already exists in the small root heap structure queue.

[0085] In S302, the first node determines dimension D iWhether it already exists in the small root heap structure queue. If so, execute S303; if not, execute S310.

[0086] S303. Take the element at the end of the heap (D tail , t tail ).

[0087] In S303, the dimension D tail is the element at the end of the heap, and t tail is the update time corresponding to the dimension D tail .

[0088] S304. Swap the element at the end of the heap D tail with the original position of D i .

[0089] S305. Determine whether t tail is less than t old .

[0090] In S305, told is the update time corresponding to the dimension D i before the update occurs. If t tail is less than t old , then execute S306; otherwise, execute S308.

[0091] S306. Perform a floating operation on (D tail , t tail ).

[0092] S307. Modify the update time corresponding to the current element at the end of the heap to t new .

[0093] In S307, the current element at the end of the heap is the dimension D i .

[0094] S308. Determine whether t tail is greater than t old .

[0095] In S308, if t tail is greater than t old , then execute S309; otherwise, it means t tail is equal to t old , and execute S307.

[0096] S309. Perform a sinking operation on (D tail , t tail ).

[0097] S310. Place (D i , t new ) at the end of the heap.

[0098] In S310, if the updated dimension Di If it does not exist in the min-heap structure queue, then (D i , t new ) is placed at the end of the heap, which is equivalent to directly adding dimension D to the min-heap structure queue i and its corresponding update time.

[0099] S311. Perform a floating-up operation on the current last element of the heap.

[0100] In S311, when a new last element of the heap is added through S310 or the update time corresponding to the current last element of the heap is modified through S307, a floating-up operation is performed on the current last element of the heap to maintain the element order of the min-heap structure queue.

[0101] It can be understood that by using a dedicated preset queue, such as a min-heap structure queue, to record the update time corresponding to the first vector clock in each dimension, the low space occupancy of the vector clock is further ensured. And through the min-heap structure queue, the first node can quickly obtain the oldest dimension information from the first queue element when it is necessary to determine the time difference, thereby further improving the processing efficiency and performance of the distributed database system.

[0102] In the embodiments of the present application, the above pruning process can effectively control the length of the vector clock and ensure that the space occupied by the vector clock is not too large. Taking the preset clock length threshold as 10 as an example, the applicant experimentally measured the space occupancy of the vector clock before and after pruning to reflect the control effect of the pruning process of the embodiments of the present application on the space occupancy of the vector clock. During the experiment, when the clock length is greater than or equal to the preset clock length threshold and the above time difference is greater than or equal to the preset time difference threshold, a pruning operation will be performed. The applicant respectively measured the space occupancy of the vector clock before and after pruning, and the effect comparison can be as Figure 9 shown.

[0103] It can be understood that the pruning operation can effectively control the space size occupied by the vector clock, pruning the redundant old dimension information in the vector clock and retaining the newer dimension information. At the same time, since removing the information on the vector clock dimension will to some extent reduce the reliability of the vector clock, that is, the timeliness of the data version cannot be guaranteed to be true. In order not to sacrifice too much the reliability of the vector clock in exchange for a decrease in space overhead, when the dimension information to be pruned from the vector clock is relatively close to the time of its last update (i.e., the time difference is less than the preset time threshold), the pruning algorithm of the embodiments of the present application will not remove it, but allows it to temporarily exceed the preset length threshold, and then remove it when these dimensions have not been updated for a long time (the time difference is greater than or equal to the preset time threshold). Experiments show that the pruning algorithm of the embodiments of the present application can reduce the space occupancy of the vector clock by about 20%.

[0104] In some embodiments, the database processing method provided by the embodiments of the present application may also be as follows Figure 10 As shown, before S101, S001 - S003 may also be executed as follows:

[0105] S001. Receive the updated data sent by the second node in the distributed system based on the data update operation; the updated data corresponds to the second vector clock.

[0106] In the embodiments of the present application, the second node is any node in the distributed system other than the first node. Data copies corresponding to the preset database file are respectively maintained on the first node and the second node. Among them, the data copy corresponding to the preset database file on the first node is the first data file, and the data copy corresponding to the preset database file on the second node is the second data file. When a data update operation occurs in the second data file of the second node, such as when a user of the second node modifies the local data copy of the second node, the second node sends the updated data to the first node based on the data update operation to synchronize the current latest data version to the first node.

[0107] In the embodiments of the present application, the updated data is the data of the changed part corresponding to the data update operation in the second data file corresponding to the second node. A second vector clock is maintained on the second node, and when the second node synchronizes data with the first node, the second vector clock is sent to the first node corresponding to the updated data.

[0108] In practical applications, database files are usually very large. When synchronizing data between nodes, the time and bandwidth required to send the entire file at one time will be very high. However, the embodiments of the present application only transmit the data change amount, which can greatly shorten the transmission time, reduce the occupation of network bandwidth, and at the same time can also reduce the burden on the server.

[0109] In some embodiments, the updated data received by the first node is encapsulated in a preset data structure.

[0110] Wherein, the preset data structure at least includes:

[0111] The operation type corresponding to the data update operation, the database name corresponding to the data update operation, the data table name corresponding to the data update operation, and at least one of the following:

[0112] The primary key value corresponding to the data update operation, the number of attributes corresponding to the primary key value, the list of names of the attributes corresponding to the data update operation, the list of types of each attribute in at least one attribute corresponding to the data update operation, and the value of each attribute in at least one attribute corresponding to the data update operation.

[0113] Exemplarily, the preset data structure corresponding to the updated data can be implemented as an Action class. The data members of the Action class can be as shown in Table 1 below:

[0114] Table 1

[0115]

[0116] Among them, op represents the operation type of the data update operation. Exemplarily, the operation type can include at least one of: insert, update, delete, and select.

[0117] schema represents the name of the database on which the data update operation acts. schema in Table 1 is of the String type.

[0118] className represents the name of the data table on which the data update operation acts. className in Table 1 is of the String type.

[0119] key represents the primary key value of the data actually operated on in the data table on which the data update operation acts. key in Table 1 is of the long type.

[0120] attrNum represents the number of attributes of the row data on which the data update operation acts. attrNum in Table 1 is of the int type.

[0121] attrName represents the list of names of the attributes on which the data update operation acts. attrName in Table 1 is an array of String type (String[]).

[0122] attrType represents the list of types of each attribute among the attributes on which the data update operation acts. attrType in Table 1 is an array of String type (String[]).

[0123] value represents the value of each attribute in at least one attribute corresponding to the data update operation, that is, the value of each attribute after the data update operation actually operates. value in Table 1 is an array of String type (String[]).

[0124] It should be noted that for different operation types, the values of key, attrNum, attrName, and attrType are also different. For some specific operation types, such as delete, one or more of the values of key, attrNum, attrName, and attrType can be empty, and specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.

[0125] It can be seen that the internal members of the Action class store the incremental data of the data update operation, as well as information such as where the incremental data is applied. In actual data transmission, the update data received by the first node can be an encapsulated Action object. When the first node uses the update data to update the first data file, accessing the Action object can obtain various necessary data information required to update the first data file.

[0126] Correspondingly, the first node can also generate update data based on the data update operations that occur on itself and send it to the second node. In some embodiments, the first node can generate each data member in the preset data structure by accessing an external Structured Query Language (SQL) command parser, thereby generating update data for the preset data structure, such as generating a corresponding Action object. That is to say, the preset data structure, such as the Action class, can be used as a data interface provided externally and applied to each node in the distributed system.

[0127] S002. Determine whether to update the first data file corresponding to the first node using the update data according to the second vector clock.

[0128] In the embodiments of the present application, the second vector clock includes second vector clock values in at least two dimensions. The first node compares at least two dimensions of the first vector clock values in its own first vector clock with at least two dimensions of the second vector clock values in the second vector clock. For the same dimension, it is determined whether the second vector clock value is greater than the first vector clock value. In the case where the second vector clock values in at least one dimension are greater than the first vector clock values, it indicates that the data version of the second node in the at least one dimension is newer than the data version of the first node, and the first node determines to update the first data file corresponding to the first node using the update data.

[0129] S003. Update the first vector clock according to the second vector clock when it is determined to update the first data file.

[0130] In the embodiments of the present application, when it is determined to update the first data file, the first node updates the first data file according to the finer data and updates the first vector clock according to the second vector clock, completing one data synchronization with the second node.

[0131] It can be understood that in the embodiments of the present application, when communicating and transmitting between nodes, only the difference data of the data operation is transmitted, and it is not necessary to transmit the entire database file, thereby reducing the amount of data and network load during communication, reducing the latency of data synchronization in the distributed system, and improving the data transmission efficiency.

[0132] The embodiment of the present application further provides a database processing device, which is applied to the first node of the embodiment of the present application. Figure 11 It is a schematic structural diagram of the database processing device provided by the embodiment of the present application. As Figure 11 shown, the database processing device 1 includes: a determination module 11 and a pruning module 12, where:

[0133] The determination module 11 is configured to determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock when the first vector clock corresponding to the first node is updated; the first vector clock includes first vector clock values of at least two dimensions; the earliest update time is the update time corresponding to the dimension that is updated earliest among the at least two dimensions;

[0134] The pruning module 12 is configured to remove the first vector clock value of the earliest updated dimension from the first vector clock values of the at least two dimensions to complete the pruning of the first vector clock when it is determined to prune the first vector clock.

[0135] In some embodiments, the determination module 11 is further configured to determine to prune the first vector clock when the time difference is greater than or equal to a preset time difference threshold; and determine not to prune the first vector clock when the time difference is less than the preset time difference threshold.

[0136] In some embodiments, the determination module 11 is further configured to determine to prune the first vector clock when the clock length is greater than or equal to a preset clock length threshold; and determine not to prune the first vector clock when the clock length is less than the preset clock length threshold.

[0137] In some embodiments, the determination module 11 is further configured to determine whether the time difference is greater than or equal to a preset time difference threshold when the clock length is greater than or equal to the preset clock length threshold; determine to prune the first vector clock when the clock length is greater than or equal to the preset clock length threshold and the time difference is greater than or equal to the preset time difference threshold; and determine not to prune the first vector clock when the clock length is less than the preset clock length threshold, or the time difference is less than the preset time difference threshold.

[0138] In some embodiments, the determination module 11 is further configured to obtain the earliest update time from a preset queue; the preset queue is used to record the time of the last update of the first vector clock value corresponding to each dimension among the at least two dimensions.

[0139] In some embodiments, the preset queue is a min-heap structure queue; the determining module 11 is further configured to obtain the earliest update time from the first queue element of the min-heap structure queue.

[0140] In some embodiments, the database processing device 1 further includes: an updating module, configured to, when it is determined that pruning is to be performed on the first vector clock, remove the current first queue element from the min-heap structure queue to update the min-heap structure queue.

[0141] In some embodiments, the updating module is further configured to, when the first vector clock corresponding to the first node is updated, determine the dimension in which the update occurs; update the update time corresponding to the dimension in which the update occurs in the min-heap structure queue according to the current time, and update the order of the min-heap structure queue.

[0142] In some embodiments, the database processing device 1 further includes: a receiving module, configured to receive update data sent by a second node in the distributed system based on a data update operation; the update data is the data of the changed part corresponding to the data update operation in a second data file corresponding to the second node; the update data corresponds to a second vector clock; the second data file is a data copy of a preset database file corresponding to the second node;

[0143] The determining module 11 is further configured to determine whether to update the first data file corresponding to the first node by using the update data according to the second vector clock; the first data file is a data copy of the preset database file corresponding to the first node;

[0144] The updating module is further configured to, when it is determined to update the first data file, update the first vector clock according to the second vector clock.

[0145] In some embodiments, the update data is encapsulated in a preset data structure; the preset data structure at least includes:

[0146] The operation type corresponding to the data update operation, the database name corresponding to the data update operation, the data table name corresponding to the data update operation, and at least one of the following:

[0147] The primary key value corresponding to the data update operation, the number of attributes corresponding to the primary key value, the list of names of the attributes corresponding to the data update operation, the list of types of each of the at least one attribute corresponding to the data update operation, and the value of each of the at least one attribute corresponding to the data update operation.

[0148] It should be noted that the description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0149] The embodiments of the present application also provide a first node. Figure 12 FIG. is an optional structural schematic diagram of the first node provided by the embodiments of the present application. As Figure 12 shown, the first node 3 includes: a memory 32 and a processor 33. Among them, the memory 32 and the processor 33 are connected through a communication bus 34; the memory 32 is used to store executable instructions; the processor 33 is used to implement the database processing method provided by the embodiments of the present application when executing the executable instructions stored in the memory 32.

[0150] The embodiments of the present application provide a readable storage medium (i.e., a computer-readable storage medium) storing executable instructions, where the executable instructions, when executed by the above-mentioned processor, will cause the above-mentioned processor to execute the database processing method provided by the embodiments of the present application.

[0151] In some embodiments, the computer-readable storage medium (i.e., the readable storage medium) may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0152] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0153] As an example, the executable instructions may or may not correspond to files in the file system, may be stored as part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0154] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected by a communication network.

[0155] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0156] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0159] As described above, only the preferred embodiments of the present application are given, and they are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A database processing method, characterized in that, applied to a first node in a distributed system, the method includes: when the first vector clock corresponding to the first node is updated, based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock, determine whether to prune the first vector clock; the first vector clock includes first vector clock values of at least two dimensions; the earliest update time is the update time corresponding to the dimension that is updated earliest among the at least two dimensions; when it is determined to prune the first vector clock, remove the first vector clock value of the earliest updated dimension from the first vector clock values of the at least two dimensions to complete the pruning of the first vector clock.

2. The method according to claim 1, characterized in that, determining whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time includes: when the time difference is greater than or equal to a preset time difference threshold, determine to prune the first vector clock; when the time difference is less than the preset time difference threshold, determine not to prune the first vector clock.

3. The method according to claim 1, characterized in that, determining whether to prune the first vector clock based on the clock length of the first vector clock includes: when the clock length is greater than or equal to a preset clock length threshold, determine to prune the first vector clock; when the clock length is less than the preset clock length threshold, determine not to prune the first vector clock.

4. The method according to claim 1, characterized in that, determining whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and the clock length of the first vector clock includes: when the clock length is greater than or equal to a preset clock length threshold, determine whether the time difference is greater than or equal to a preset time difference threshold; when the clock length is greater than or equal to the preset clock length threshold and the time difference is greater than or equal to the preset time difference threshold, determine to prune the first vector clock; when the clock length is less than the preset clock length threshold, or the time difference is less than the preset time difference threshold, determine not to prune the first vector clock.

5. The method according to claim 2 or 4, characterized in that, the method further includes: obtain the earliest update time from a preset queue; the preset queue is used to record the time of the last update of the first vector clock value corresponding to each dimension among the at least two dimensions.

6. The method according to claim 5, characterized in that, the preset queue is a min-heap structure queue; obtaining the earliest update time from the preset queue includes: obtain the earliest update time from the dimension corresponding to the first queue element of the min-heap structure queue.

7. The method according to claim 6, wherein, the method further comprises: when it is determined to prune the first vector clock, removing the dimension corresponding to the current first queue element from the small root heap structure queue to update the small root heap structure queue.

8. The method according to claim 6, wherein, the method further comprises: when the first vector clock corresponding to the first node is updated, determining the updated dimension; according to the current time, updating the update time corresponding to the updated dimension in the small root heap structure queue, and updating the order of the small root heap structure queue.

9. The method according to any one of claims 1-4, or any one of claims 6-8, wherein, the method further comprises: receiving update data sent by a second node in the distributed system based on a data update operation; the update data is the data of the changed part corresponding to the data update operation in a second data file corresponding to the second node; the update data corresponds to a second vector clock; the second data file is a data copy of a preset database file corresponding to the second node; determining whether to update a first data file corresponding to the first node according to the second vector clock; the first data file is a data copy of the preset database file corresponding to the first node; when it is determined to update the first data file, updating the first vector clock according to the second vector clock.

10. The method according to claim 9, wherein, the update data is encapsulated in a preset data structure; the preset data structure at least includes: the operation type corresponding to the data update operation, the database name corresponding to the data update operation, the data table name corresponding to the data update operation, and at least one of the following: the primary key value corresponding to the data update operation, the number of attributes corresponding to the primary key value, the list of names of attributes corresponding to the data update operation, the list of types of each of at least one attribute corresponding to the data update operation, and the value of each of at least one attribute corresponding to the data update operation.

11. A database processing device, wherein, applied to a first node in a distributed system, the device comprises: a determination module, configured to determine whether to prune the first vector clock based on the time difference between the earliest update time corresponding to the first vector clock and the current time, and / or the clock length of the first vector clock when the first vector clock corresponding to the first node is updated; the first vector clock includes first vector clock values of at least two dimensions; the earliest update time is the update time corresponding to the earliest updated dimension among the at least two dimensions; a pruning module, configured to, when it is determined to prune the first vector clock, remove the first vector clock value of the earliest updated dimension from the first vector clock values of the at least two dimensions to complete the pruning of the first vector clock.

12. A first node, wherein, Comprising: A memory and a processor; wherein, The memory is used for storing executable instructions; The processor is used for implementing the method according to any one of claims 1-10 when executing the executable instructions stored in the memory.

13. A readable storage medium, Characterized in that It stores executable instructions for causing a processor to implement the method according to any one of claims 1 to 10 when executed.

Citation Information

Cited By

  • Data query method, system and equipment based on vector retrieval acceleration and medium

    CN120849411A