Distributed database architecture system, metadata node and method thereof
By introducing metadata nodes and multiple reference mechanisms into the distributed database architecture system, the problem of binding between computing nodes and storage nodes in the existing technology is solved, flexible provisioning and rapid upgrades are achieved, and low-latency access needs are met, and the performance and availability of the database are improved.
Patent Information
- Application Number
- CN202311489457.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-09
AI Technical Summary
The existing distributed database architecture cannot realize the flexible provisioning of computing nodes and storage nodes and the rapid upgrade of database systems, and it is also difficult to achieve low-latency data access.
By introducing metadata nodes into the distributed database architecture system, recording the correspondence between the data blocks of each data stream and the storage location, and creating multiple references for each data stream, allowing different computing nodes to independently access the same data stream, thereby realizing the unbinding and flexible scheduling of the computing nodes and storage nodes.
It realizes flexible provisioning of computing nodes and storage nodes and rapid upgrade of database systems, while meeting the low-latency data access requirements and improving the performance, availability and flexibility of the database.
Smart Images

Figure CN119961349A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of distributed databases, and more specifically, to a distributed database architecture system, a metadata node for a distributed database architecture system, a method for a distributed database architecture system, and a method for a metadata node for a distributed database architecture system. Background Art
[0002] With the rapid development of the Internet, the growth of data volume and the number of users, data analysis and processing are becoming increasingly important. The existing distributed database architecture still cannot achieve low latency, efficient access, and fast upgrades in many scenarios, and there is still room for improvement. Summary of the invention
[0003] In view of the above, the present disclosure provides a distributed database architecture system, which can achieve flexible deployment of computing nodes and storage nodes and rapid upgrade of the database system while meeting low-latency data access.
[0004] According to one aspect of the present disclosure, a distributed database architecture system is provided, comprising a plurality of computing nodes, a plurality of storage nodes and metadata nodes, wherein the metadata nodes are configured to: record the correspondence between each data block in the one or more data blocks into which each data stream is divided and the storage location for storing the each data block, the storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes, and record more than two references to a first data stream to be read by more than two computing nodes, each reference is used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream. The plurality of computing nodes are configured to: read each data stream from the plurality of storage nodes or store each data stream to the plurality of storage nodes according to the correspondence for each data stream; and the plurality of storage nodes are configured to: store the one or more data blocks of each data stream in the storage location for the one or more data blocks, respectively.
[0005] Optionally, the metadata node is further configured to: in response to a first computing node among the more than two computing nodes needing to read the first data stream, generate a reference to the first data stream for the first computing node; in response to the first computing node no longer needing to read the first data stream, delete the reference to the first data stream for the first computing node; and in response to all references to the first data stream being deleted, instruct the multiple storage nodes to delete the first data stream.
[0006] Optionally, the metadata node is further configured to: in response to a storage request from a second computing node among the multiple computing nodes to store the first data stream, allocate one or more first storage locations in one or more first storage nodes among the multiple storage nodes to the one or more first data blocks into which the first data stream is divided, return information about the one or more first storage locations for storing the one or more first data blocks to the second computing node, and record the correspondence between each first data block in the one or more first data blocks and the first storage location for storing each first data block. The second computing node is configured to: divide the first data stream into the one or more first data blocks, and send the corresponding first data block in the one or more first data blocks to the first storage node corresponding to each first storage location according to the information about the one or more first storage locations. The first storage node is configured to: store the first data block received from the second computing node in the corresponding first storage location.
[0007] Optionally, the metadata node is also configured to: record information of the newly added storage node, and in response to a request from a third computing node among the multiple computing nodes to store a data stream, allocate a storage location including a storage location located in the newly added storage node to the third computing node.
[0008] Optionally, the distributed database architecture system also includes a control node, wherein the control node is configured to: send information of the newly added storage node to the metadata node based on the newly added storage node.
[0009] Optionally, the distributed database architecture system also includes a control node, wherein the control node is configured to: schedule a second data stream processed by a fourth computing node among the multiple computing nodes to be processed by a fifth computing node among the multiple computing nodes according to the load distribution of the multiple computing nodes, and control the metadata node to establish a reference to the second data stream for the fifth computing node during the scheduling, and delete the reference to the second data stream for the fourth computing node after the scheduling is completed.
[0010] Optionally, the distributed database architecture system also includes a control node, wherein the control node is configured to: control the metadata node to copy the reference to the data stream in the first data stream set for each computing node in the first computing node group for processing the first data stream set to create a reference to the data stream in the first data stream set for each computing node in the newly added or original second computing node group.
[0011] Optionally, the control node is further configured to: control the first data stream set to be offline from the first computing node group; and the metadata node is further configured to: delete references to data streams in the first data stream set for each computing node in the first computing node group.
[0012] Optionally, the control node is further configured to: control each computing node in the first computing node group to synchronize data in a memory to a memory of a computing node in the second computing node group.
[0013] Optionally, the control node is further configured to: when the difference between the data in the memory of the first computing node group and the data in the memory of the second computing node group is less than a first predetermined threshold or the synchronization delay between the first computing node group and the second computing node group is lower than a second predetermined threshold, control the first computing node group to stop writing data into its memory; after the synchronization is completed, control the second computing node group to go online, and control the first computing node group to go offline; and after the synchronization is completed, control the metadata node to delete the reference to the data stream in the first data stream set for each computing node of the first computing node group.
[0014] Optionally, the metadata node is also configured to store multiple copies for each reference.
[0015] Optionally, the multiple computing nodes are deployed in multiple cloud services; and / or the multiple storage nodes are deployed in multiple cloud services.
[0016] According to another aspect of the present disclosure, a metadata node for a distributed database architecture system is provided, the distributed database architecture system comprising a plurality of computing nodes, a plurality of storage nodes and the metadata node, wherein the metadata node is configured to: record a correspondence between each data block in one or more data blocks into which each data stream is divided and a storage location for storing the each data block, the storage location of the one or more data blocks for each data stream being located in one or more of the plurality of storage nodes; record more than two references to a first data stream to be read by more than two computing nodes, each reference being used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream.
[0017] According to another aspect of the present disclosure, a method for a distributed database architecture system is provided, wherein the distributed database architecture system includes multiple computing nodes, multiple storage nodes and metadata nodes, wherein the method includes: the metadata node: recording the correspondence between each data block in the one or more data blocks into which each data stream is divided and the storage location for storing the each data block, the storage location of the one or more data blocks for each data stream is located in one or more of the multiple storage nodes, and recording more than two references to a first data stream to be read by more than two computing nodes, each reference is used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream; the multiple computing nodes: reading each data stream from the multiple storage nodes or storing each data stream to the multiple storage nodes according to the correspondence for each data stream; and the multiple storage nodes: storing the one or more data blocks of each data stream in the storage location resources for the one or more data blocks, respectively.
[0018] According to another aspect of the present disclosure, a method for a metadata node of a distributed database architecture system is provided, wherein the distributed database architecture system includes multiple computing nodes, multiple storage nodes and the metadata node, wherein the method includes: recording a correspondence between each data block in one or more data blocks into which each data stream is divided and a storage location for storing the each data block, wherein the storage location of the one or more data blocks for each data stream is located in one or more of the multiple storage nodes; and recording more than two references to a first data stream to be read by more than two computing nodes, wherein each reference is used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream.
[0019] According to the embodiments of the present disclosure, the data blocks of each data stream are allocated storage locations that can be located in different storage nodes through metadata nodes, which realizes the unbinding and flexible scheduling of computing nodes and storage nodes, and the corresponding relationship between each data block of each data stream and the storage location storing the data block is recorded through metadata nodes, so that the stored data stream can be quickly read. In addition, by creating multiple references to a single data stream for different computing nodes, different computing nodes can independently access the same data stream, thereby realizing flexible deployment and rapid upgrade of the database, and managing the life cycle of the data stream through multiple references. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 An existing distributed database architecture is shown.
[0022] Figure 2 A structural diagram of a distributed database architecture system according to an embodiment of the present disclosure is shown.
[0023] Figure 3 A schematic diagram is shown of a storage node storing a data stream in a storage location of each data block in a data block manner according to an embodiment of the present disclosure.
[0024] Figure 4 A schematic diagram showing the relationship between computing nodes, metadata nodes, and storage nodes according to an embodiment of the present disclosure.
[0025] Figure 5 A schematic diagram showing a process of writing a data stream according to an embodiment of the present disclosure is shown.
[0026] Figure 6 A schematic diagram is shown of how to quickly perform reading and writing after a storage node is added according to an embodiment of the present disclosure.
[0027] Figure 7 A schematic diagram of a database fast zero-copy multi-copy architecture according to an embodiment of the present disclosure is shown.
[0028] Figure 8 A schematic diagram of a smooth upgrade process based on fast database cloning according to an embodiment of the present disclosure is shown.
[0029] Fig. 9 A flowchart of a method for a distributed database architecture system according to an embodiment of the present disclosure is shown.
[0030] Fig.10 A flowchart of a method for a metadata node in a distributed database architecture system according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] Reference will now be made in detail to specific embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Although the present disclosure will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit the present disclosure to the described embodiments. On the contrary, it is intended to cover changes, modifications and equivalents included within the spirit and scope of the present disclosure as defined by the appended claims. It should be noted that the method steps described herein can all be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of the two.
[0032] The current distributed database architecture binds computing nodes and storage nodes, which makes it impossible to flexibly deploy computing nodes and storage nodes and quickly upgrade the database system. For example, when changing or expanding computing nodes or storage nodes, a large amount of data needs to be moved, which takes a long time. For example, a common database uses a key-value method to store data, where a key is associated with a value as a unique identifier. When querying, the data is found by finding the storage location of the data through the key. However, when this database encounters the expansion of storage nodes, it cannot immediately provide write services to the storage nodes because keys need to be assigned to the data.
[0033] Figure 1 FIG. 1 shows an existing distributed database architecture system 100. Figure 1 As shown, the distributed database architecture system 100 includes computing nodes 110 and storage nodes 120. Each storage node 120 corresponds to a computing node 110, and the computing node 110 controls the reading and writing of data in the storage node 120. When a storage node 121 is added to the computing node 110, or the storage node 120 bound to the computing node 110 is changed to the storage node 121, the data writing service cannot be immediately provided in the storage node 121. It is necessary to wait until part of the database data is moved from the existing storage node 120 to the storage node 121 before the data writing service can be provided in the storage node 121, which takes a long time. In addition, the existing object storage distributed database architecture cannot achieve low-latency access.
[0034] In view of the above, the present disclosure proposes a new distributed database architecture system, which can realize flexible deployment of computing nodes and storage nodes and rapid upgrade of database systems, while meeting low-latency data access. The distributed database architecture system of the present disclosure is suitable for scenarios that require processing large amounts of data and low-latency access, and provides users with high-performance, high-availability and flexible database services. The distributed database architecture system of the present disclosure can be implemented based on standard cloud service products provided by one or more cloud service providers.
[0035] Figure 2A structural diagram of a distributed database architecture system 200 according to an embodiment of the present disclosure is shown.
[0036] The distributed database architecture system 200 may include multiple computing nodes 210, multiple storage nodes 220, and metadata nodes 230. As an example, Figure 2 Only one computing node and one storage node are shown. Optionally, the distributed database architecture system 200 may also include a control node 240. The control node 240 may be used to apply for cloud service resources from a cloud service provider to construct computing nodes 210, storage nodes 220, and metadata nodes 230, and may perform a series of controls, such as monitoring loads, scaling nodes, upgrading and downgrading databases, and the like.
[0037] like Figure 2 As shown in , computing node 210, storage node 220 and metadata node 230 can be deployed in cloud service resources provided by cloud service providers. For example, computing node 210 can be a computing resource allocated on cloud service, which can be provided by any hardware with processing capability, such as a central processing unit (CPU). Storage node 220 can be a persistent storage resource allocated on cloud service, which can be provided by any hardware with persistent storage capability, such as a magnetic hard disk and / or a solid state drive (SSD). Metadata node 230 can be a resource that provides metadata storage and processing on cloud service, for example, it can be provided by hardware with processing and storage capabilities. Metadata node 230 can also be implemented by computing node and storage node. Control node 240 can be any computing resource (such as computing resource allocated on cloud service), which can be provided by any hardware with processing capability, such as a central processing unit (CPU). Control node 240 can also be implemented on computing node. The hardware providing resources can provide services in conjunction with corresponding software. According to an embodiment of the present disclosure, multiple computing nodes can be deployed in one or more cloud services. Multiple storage nodes can also be deployed in one or more cloud services.
[0038] The metadata node 230 may be configured to record the correspondence between each data block in the one or more data blocks into which each data stream is divided and the storage location for storing each data block. The storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes.
[0039] In an embodiment of the present disclosure, a data stream refers to a string of data that is streamed for transmission or storage. A data stream can be a stream in any form, such as a character stream, a byte stream, a bit stream, etc. For example, in accessing a database, the data in the database can usually be converted into a data stream, such as a byte stream. A data stream can represent a data stream corresponding to data stored in a stream. For example, a data stream can correspond to a data stream corresponding to a record in a database (such as a data table). In an embodiment of the present disclosure, a database can be a collection of data streams.
[0040] The data stream may be divided into one or more data blocks so as to be stored in one or more storage locations in one or more storage nodes in a dispersed manner in units of data blocks. Generally, a data block is a segment of a data stream, and a data stream has multiple data blocks. In an embodiment of the present disclosure, the number of data blocks of a data stream may be arbitrary and has no upper limit. In the case where the data stream is relatively small, a data stream may also be divided into only one data block. The size of the data block may be determined according to a specific application scenario. The metadata node, the computing node, and the storage node are all processed according to the determined data block size. Each data block may be stored in a storage location. In an embodiment of the present disclosure, different data blocks may be stored in the storage locations of the same or different storage nodes. In the case where a data stream is divided into multiple data blocks, different data blocks of a data stream may be stored in the storage locations of the same or different data nodes. Therefore, in an embodiment of the present disclosure, a data stream may be stored in different storage nodes, thereby realizing the dispersed and flexible storage of the data stream. Moreover, the computing node processing the data stream may access multiple storage nodes, and the same storage node may also be accessed by multiple computing nodes, eliminating the binding relationship between the computing node and the storage node. In addition, the data stream is streamed data, so the access to the data stream and the data blocks it is divided into does not depend on the key-value relationship. Therefore, when the data blocks of the data stream need to be stored in a new storage node, they can be stored directly without first copying part of the data from other data nodes to the new data node or creating a key-value relationship in the new data node, thereby achieving rapid expansion of storage nodes.
[0041] In order to facilitate the rapid reading of the stored data stream, the metadata node 230 can record the corresponding relationship between each data block in one or more data blocks into which each data stream is divided and the storage location for storing each data block as one of the metadata. In other words, the metadata node 230 records the storage location of the database for each data block in units of data blocks, that is, the metadata according to the embodiment of the present disclosure is metadata at the data block level. For example, the corresponding relationship can be a mapping relationship between the identifier of the data block in the data stream and the identifier of the storage location of the data block. The storage location according to the embodiment of the present disclosure can have different granularities, such as storage node granularity, storage disk granularity in the storage node, or more specific storage location granularity in the storage disk. In the case of using storage node granularity, the metadata node only allocates and records the storage node storing a certain data block, and which storage disk of the storage node is specifically stored and which specific location in the storage disk can be determined by the storage node. In this case, the identification of the above storage location can only include the identification of the storage node. In the case of adopting the storage disk granularity in the storage node, the metadata node allocates and records the storage node storing a certain data block and the storage disk storing the data block in the storage node, and the specific location in the storage disk where the data block is stored can be determined by the storage node. In this case, the identification of the above storage location may include the identification of the storage node and the identification of the storage disk in the storage node. In the case of adopting a more specific storage location granularity in the storage disk, the identification of the location needs to include more specific storage location information in addition to the identification of the storage node and the identification of the storage disk in the storage node. Based on the above correspondence, the location where a certain data stream is stored can be quickly found to achieve fast data reading. The implementation method of the correspondence is not limited to the above specific embodiment, as long as the above purpose is achieved.
[0042] The following Tables 1 and 2 show two examples of the above correspondence: Table 1 shows an example in which the storage location is at the storage node granularity, and Table 2 shows an example in which the storage location is at the storage disk granularity.
[0043] Table 1
[0044] Data flow identification Data block identifier Storage Location 1 1 Storage Node 1 1 2 Storage Node 1 1 3 Storage Node 2 2 1 Storage Node 1 2 2 Storage Node 2
[0045] As shown in Table 1, data stream 1 is divided into three data blocks 1, 2, and 3, wherein data blocks 1 and 2 are stored in storage node 1, and data block 3 is stored in storage node 2. Therefore, the corresponding relationship may include that data block identifiers 1, 2, and 3 corresponding to data stream identifier 1 correspond to storage nodes 1, 1, and 2, respectively. In addition, data stream 2 is divided into two data blocks 1 and 2, wherein data block 1 is stored in storage node 1, and data block 2 is stored in storage node 2. Therefore, the corresponding relationship may include that data block identifiers 1 and 2 corresponding to data stream identifier 2 correspond to storage nodes 1 and 2, respectively.
[0046] Table 2
[0047] Data flow identification Data block identifier Storage Location 1 1 Disk C(1-C) in storage node 1 1 2 Disk D(1-D) in storage node 1 1 3 Disk C (2-C) in storage node 2 2 1 Disk D(1-D) in storage node 1 2 2 Disk C (2-C) in storage node 2
[0048] As shown in Table 2, data stream 1 is divided into three data blocks 1, 2, and 3, wherein data blocks 1 and 2 are respectively stored in disks C and D of storage node 1 (for example, they can be identified as 1-C and 1-D, respectively), and data block 3 is stored in disk C of storage node 2 (for example, they can be identified as 2-C). Therefore, the corresponding relationship may include that data block identifiers 1, 2, and 3 corresponding to data stream identifier 1 correspond to storage location identifiers 1-C, 1-D, and 2-C, respectively. In addition, data stream 2 is divided into two data blocks 1 and 2, wherein data block 1 is stored in disk D of storage node 1 (for example, they can be identified as 3-D), and data block 2 is stored in disk C of storage node 2 (for example, they can be identified as 4-C). Therefore, the corresponding relationship may include that data block identifiers 1 and 2 corresponding to data stream identifier 2 correspond to storage locations 3-D and 4-C, respectively.
[0049] It should be noted that the above Tables 1 and 2 are only examples of corresponding relationships, and the present disclosure is not limited thereto, and the corresponding relationships may also be in other forms. For example, the data stream identifiers and data block identifiers in the above Tables 1 and 2 may be combined, for example, data block 1 of data stream 1 may be represented as data block 1-1, data block 2 of data stream 1 may be represented as data block 1-2, data block 3 of data stream 1 may be represented as data block 1-3, data block 1 of data stream 2 may be represented as data block 2-1, and data block 2 of data stream 2 may be represented as data block 2-2.
[0050] Optionally, this correspondence and other metadata in the metadata node can be replicated into multiple copies to ensure the reliability of the metadata. When a copy is damaged, the metadata can be provided through other copies. In addition, the metadata node can also allocate one or more copies of each data block to matching storage nodes according to the data redundancy strategy. In the case of missing copies, the metadata node can initiate a data block replication request to the storage node to meet the data redundancy strategy.
[0051] Multiple computing nodes 210 can be configured to read each data stream from the multiple storage nodes or store each data stream to the multiple storage nodes according to the corresponding relationship for each data stream. According to an embodiment of the present disclosure, since the correspondence between one or more data blocks into which each data stream is divided and the storage locations storing these data blocks are recorded by metadata nodes, and multiple computing nodes can access the corresponding relationship in the metadata nodes, multiple computing nodes can access these data blocks of the data stream in multiple storage nodes according to the corresponding relationship. Therefore, the binding relationship between computing nodes and storage nodes is eliminated, a computing node can access multiple storage nodes as needed, and the same storage node can also be accessed by multiple computing nodes as needed. This can be achieved by allocating storage locations to each data block of the data stream through metadata nodes.
[0052] For storage operations, the metadata node 230 can allocate one or more storage locations for one or more first data blocks into which the data stream is divided based on a request from a computing node to store the data stream, and return information about the allocated storage locations to the computing node. Thus, the computing node can send the corresponding data block to the corresponding storage node based on the information of each storage location so that the database is stored in the corresponding storage location. It can be seen that the computing node can store the corresponding data block in the corresponding data node according to the storage node assigned by the metadata node, so that it is not limited to storing in a fixed storage node. The metadata node can flexibly allocate storage locations based on information such as the storage nodes in the system and their storage status. It should be noted that the identification information of the above-mentioned storage node is also included in the correspondence between the data blocks of the data stream and the storage locations, so that the computing node can determine the corresponding storage node based on the correspondence.
[0053] For read operations, if a computing node expects to read a data stream, it can obtain the correspondence between the data blocks and storage locations of the data stream from the metadata node, and then read each data block of the data stream from the corresponding storage location of the corresponding storage node, thereby reading the entire data stream.
[0054] Accordingly, multiple storage nodes 220 can be configured to store the one or more data blocks of each data stream in storage locations for the one or more data blocks, respectively. For example, a storage node that receives a request to store a certain data block from a computing node can store the data block in a certain storage address according to the storage location information contained in the request. For example, if the storage location information includes information of the target storage disk, the storage node stores the database in the target storage disk, and the specific storage address is determined by the storage node according to its own storage strategy; if the storage location information itself includes a more specific storage address, the storage node stores the database in the storage address contained in the storage location information; if the storage location information only has storage node information and no further information, the storage node determines the specific storage disk and address according to its own storage strategy.
[0055] Figure 3 A schematic diagram is shown of a storage node storing a data stream in a data block resource in a data block manner according to an embodiment of the present disclosure.
[0056] A storage node is a node that stores data, which may include and be responsible for managing storage media. The storage medium may be any form of persistent storage medium, such as a magnetic hard disk, a solid-state hard disk, etc. The storage node stores the data block in the storage address in the storage medium. The storage node can handle the reading and writing of data blocks, as well as the deletion and copying requests of data blocks, and verify the integrity of data blocks to ensure the persistence and correctness of the data. The storage node can support multiple computing nodes to read data at the same time, including reading the same data at the same time.
[0057] like Figure 3 As shown, the storage medium in the data node stores the data stream in the form of data blocks. A data stream can be divided into multiple data blocks and stored in multiple storage nodes. There is a corresponding relationship between the data blocks into which the data stream is divided and the storage locations of the stored data blocks. In this way, a data stream can be stored in different data nodes, realizing the dispersion and flexible storage of data streams. Moreover, the same computing node can access multiple storage nodes, and the same storage node can also be accessed by multiple computing nodes, eliminating the binding relationship between computing nodes and storage nodes, and realizing flexible scheduling between computing nodes and storage nodes.
[0058] return Figure 2, the metadata node 230 may also be configured to record more than two references to a first data stream to be read by more than two computing nodes. Each reference is used by one of the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream. According to an embodiment of the present disclosure, multiple references can be created for a data stream, each reference is used for a computing node, so that multiple references can be used for multiple different computing nodes. The reference to the data stream represents the association relationship with the data stream, for example, it can be an identifier of the data stream, and the identifier is associated with the data stream. Since the reference can be associated with the data stream, the corresponding relationship between the data block of the data stream and the storage location storing the data block of the data stream can be obtained according to the reference, and then the data stream can be stored and read. Recording multiple references to a data stream can, for example, record multiple identifiers for the data stream, each identifier is used for a computing node. The values of these identifiers may be the same or different. In the case where respective references (such as identifiers) are set for different computing nodes for the same data stream, from the perspective of the computing node, it is shown that it enjoys the data stream alone. Therefore, when a computing node needs to read a data stream, it only needs to add a reference to the computing node. Conversely, if the computing node no longer needs to read the data stream, it only needs to delete the reference to the computing node. In this way, flexible deployment and rapid upgrade of the database can be achieved.
[0059] In addition, the metadata node can also manage the life cycle of the data stream and the life cycle of a certain computing node being able to access the data stream by setting different references to different computing nodes. Specifically, the metadata node can be configured to: in response to a first computing node among the two or more computing nodes needing to read the first data stream, generate a reference to the first data stream for the first computing node; in response to the first computing node no longer needing to read the first data stream, delete the reference to the first data stream for the first computing node; and in response to all references to the first data stream being deleted, instruct the multiple storage nodes to delete the first data stream.
[0060] The life cycle of a data stream starts when the data stream is stored, that is, when the metadata node creates the first reference, and ends when the data stream is deleted, that is, when all references to the data stream are deleted. The life cycle of each computing node accessing the data stream starts when the computing node needs to access the data stream, that is, when a reference is created for the computing node, and ends when the computing node no longer needs to access the data stream, that is, when the reference for the computing node is deleted. By creating respective references for different computing nodes, the metadata node can clearly know which computing nodes may access the data stream, and can flexibly manage the life cycle of the data stream and the life cycle of each computing node accessing the data stream.
[0061] The following Table 3 shows an example of the above-mentioned citations, but the present disclosure is not limited thereto.
[0062] Table 3
[0063] References Compute Node Data Flow Compute Node 1 - Data Stream 1 Compute Node 1 Data Stream 1 Compute Node 1 - Data Flow 2 Compute Node 1 Data Flow 2 Compute Node 2 - Data Flow 1 Compute Node 2 Data Stream 1 Compute Node 2 - Data Flow 2 Compute Node 2 Data Flow 2 Compute Node 3 - Data Flow 3 Compute Node 3 Data Flow 3 Compute Node 4 - Data Flow 3 Compute Node 4 Data Flow 3
[0064] As shown in Table 3, for example, if computing nodes 1 and 2 both want to access (read) data stream 1, a reference to data stream 1 is established for computing node 1 (for example, identified as "computing node 1-data stream 1"), and a reference to data stream 1 is established for computing node 2 (for example, identified as "computing node 2-data stream 1"). In this way, computing nodes 1 and 2 can access data stream 1 at the same time. If computing nodes 1 and 2 both want to access data stream 2, a reference to data stream 2 is established for computing node 1 (for example, identified as "computing node 1-data stream 2"), and a reference to data stream 2 is established for computing node 2 (for example, identified as "computing node 2-data stream 2"). In this way, computing nodes 1 and 2 can access data stream 2 at the same time. The same is true for computing nodes 3 and 4, which will not be repeated here. Of course, the form of the reference is not limited to the form of the example in Table 3. The reference only needs to indicate which computing node is served and which data stream is used to read.
[0065] Optionally, the metadata node can also store multiple copies of each reference to meet data redundancy requirements so that reference information can be obtained from other copies when one copy is damaged.
[0066] Figure 4 An example schematic diagram showing the relationship between computing nodes, metadata nodes, and storage nodes according to an embodiment of the present disclosure is shown.
[0067] like Figure 4As shown, the metadata node 430 records the correspondence 437 between the multiple data blocks 4211, 4212, 4221 divided by the first data stream and the multiple storage locations in the storage nodes 421, 422 for storing the multiple data blocks 4211, 4212, 4221 of the first data stream, and the correspondence 438 between the second data stream and the multiple storage locations in the storage nodes 422, 423 for storing the multiple data blocks 4222, 4231, 4232 divided by the second data stream.
[0068] For example, the correspondence 437 is used to store two data blocks 4211 and 4212 of the first data stream in the storage node 421 and another data block 4221 in the storage node 422. Accordingly, the correspondence 437 can be used to read three data blocks 4211, 4212, and 4221 from the storage node 421 and the storage node 422 to completely read the first data stream. The correspondence 438 is used to store two data blocks 4231 and 4232 of the second data stream in the storage node 423 and another data block 4222 in the storage node 422. Accordingly, the correspondence 438 can be used to read three data blocks 4231, 4232, and 4222 from the storage node 423 and the storage node 422 to completely read the second data stream.
[0069] In the distributed storage service of the prior art, the relationship between the computing node and the storage node it can access is bound, and the data stream processed by a computing node is difficult to store in multiple storage nodes. If a new storage node is added, it is necessary to move the data to the newly added storage node in order to bind the original computing node to the newly added storage node, which cannot achieve rapid expansion or migration. However, according to an embodiment of the present disclosure, the computing nodes and storage nodes can achieve flexible allocation of computing nodes and storage nodes by establishing and recording the relationship between the data blocks and storage locations of the data stream in the metadata node.
[0070] In addition, if Figure 4 As shown, the metadata node 430 may also record more than two references to the first data stream to be read by more than two computing nodes as part of the metadata. Each reference is used by each of the more than two computing nodes to obtain the corresponding relationship of the first data stream to read the first data stream. The number of references to the first data stream may be the same as the number of computing nodes to access the first data stream, and one reference to the first data stream corresponds to one computing node to access the first data stream, and the corresponding relationship is dedicated to the computing node, so that each computing node can independently use the reference to access data without interfering with each other.
[0071] For example, the metadata node 430 records multiple references 431, 432, 433, 434 for the first data stream, which are used for computing nodes 411, 412, 413, 414 to obtain the corresponding relationship 437 for the first data stream so as to access each data block 4211, 4212, 4221 of the data stream. References 431, 432, 433, 434 point to the same data stream (the first data stream), but are used for computing nodes 411, 412, 413, 414, respectively. Thus, each computing node 411, 412, 413, 414 can simultaneously obtain the corresponding relationship 437 of each data block of the first data stream, so that the data blocks 4211, 4212, 4221 in the first data stream can be accessed simultaneously and independently. That is, a data stream can be accessed by multiple computing nodes at the same time. From the perspective of the computing node, multiple references are manifested as each computing node having its own data stream, and thus having its own database, that is, a data stream set consisting of one or more data streams. From the perspective of the database, multiple references are manifested as the database being replicated as multiple databases in a zero-copy manner, each for different computing nodes. Different computing nodes each have their own references rather than sharing the same reference, which can achieve flexible deployment and rapid upgrade of the database.
[0072] In addition, a computing node may also use multiple references to multiple data streams to access (i.e., read) multiple data streams. For example, if computing node 413 wants to access the first data stream and the second data stream, metadata node 430 creates reference 433 for computing node 413 to access the first data stream, and creates reference 435 to access the second data stream.
[0073] The specific forms of references can be varied. For example, a reference can be an identifier of a data stream, which is associated with the data stream. In addition, the reference can also be attached with an identifier of a computing node that uses the reference, thereby indicating which computing node uses the reference. Of course, it is also possible to implicitly indicate which computing node uses the reference, for example, a reference recorded in a specific location implicitly indicates that it is specific to a computing node, or the value of the reference itself implicitly indicates that it is specific to a computing node. In addition, the reference can also contain a correspondence between a data block of a data stream and a storage location, for example, it can be the correspondence itself, that is, a copy of the correspondence can be recorded for each computing node that wants to access the data stream, dedicated to the computing node. The operation on the reference is a pure metadata operation, does not involve the copy of data, and can be completed quickly. After establishing a reference to a data stream for a computing node, the computing node will see an independent data stream through the reference. When all references to a data stream are released (i.e., deleted), the metadata node will issue a deletion request to the storage node where the data block contained in the data stream is located, and delete the metadata of the data stream and the data block, so that the life cycle of the data stream and the life cycle of the data stream that can be accessed by a computing node can be managed by reference.
[0074] According to other embodiments of the present disclosure, a reference may also be used by multiple computing nodes to read the same data stream. In this case, if a computing node no longer uses the reference, the reference will not be deleted but will be reserved for other computing nodes to continue using. For example, in an application scenario, if a computing node has a large throughput for the data stream, other computing nodes may be temporarily established to use the reference of the computing node to share the throughput of the computing node. When other computing nodes are not needed to share, other computing nodes may be offline, but the reference of the data stream is still reserved for use by the one computing node.
[0075] According to the embodiments of the present disclosure, a storage location that can be located in different storage nodes is allocated to the data block of each data stream through a metadata node, thereby realizing the unbinding and flexible scheduling of computing nodes and storage nodes, and the corresponding relationship between each data block of each data stream and the storage location storing the data block is recorded through the metadata node, so that the stored data stream can be quickly read. In addition, by creating multiple references to a single data stream for different computing nodes, different computing nodes can independently access the same data stream, thereby realizing flexible deployment and rapid upgrade of the database.
[0076] Figure 5 A schematic diagram showing a process of writing a data stream according to an embodiment of the present disclosure is shown.
[0077] Figure 5The second computing node 511 initiates writing (storing) the first data stream as an example for description. First, the second computing node 511 sends a storage request to the metadata node 530 for storing the first data stream.
[0078] In response to the storage request of the second computing node 510 to store the first data stream, the metadata node 530 allocates one or more first storage locations in one or more first storage nodes 521, 522 in the plurality of storage nodes to the one or more first data blocks into which the first data stream is divided, and returns information of the one or more first storage locations for storing the one or more first data blocks to the second computing node 510, and records the correspondence between each first data block in the one or more first data blocks and the first storage location for storing each first data block. For example, the first data stream is divided into two data blocks, the first data block is allocated for storage in the first storage node 521, so that the identification information ID1 of the first storage node 521 can be returned for the first data block; the second data block is allocated for storage in the first storage node 522, so that the identification information ID2 of the first storage node 522 can be returned for the second data block. Of course, if the above storage location includes a more specific storage location in addition to the storage node, the information of the returned storage location can also include information of the more specific storage location, such as information of the storage disk in the storage node.
[0079] The second computing node 510 divides the first data stream into the one or more first data blocks ( Figure 5 There are two data blocks in the storage location), and according to the information of the one or more first storage locations, the corresponding first data blocks in the one or more first data blocks are sent to the first storage nodes 521 and 522 corresponding to each first storage location. For example, the storage location of the first data block is located in the first storage node 521, then when storing the first data block, the second computing node 510 sends it to the first storage node 521 according to the identification information ID1 of the first storage node 521. The storage location corresponding to the second data block is located in the first storage node 522, then when storing the second data block, the second computing node 510 sends it to the first storage node 522 according to the identification information ID2 of the first storage node 522.
[0080] In one embodiment, before the second computing node 510 sends the corresponding data block to the corresponding first storage node 521, 522, the corresponding data block can be written to the memory until the data in the memory reaches a predetermined amount of data, and the second computing node 510 then asynchronously writes the data block cached in the memory to the first storage node.
[0081] Finally, the first storage nodes 521 and 522 store the first data block received from the second computing node 510 into the corresponding first storage location.
[0082] The following describes specific implementation methods and technical effects of embodiments of the present disclosure in different scenarios.
[0083] Scenario 1: Rapid expansion of storage nodes
[0084] Over time, the amount of data stored in a database may grow rapidly. If the storage space is insufficient, the database will not be able to accommodate new data, causing the system to crash or not work properly. At the same time, when the database storage space is close to full, the database performance will be affected. Read and write operations may become slower, query response time will increase, and user experience will decrease. It can be seen that quickly expanding storage space is very important to ensure the normal operation of the system, improve performance, and protect data security.
[0085] Therefore, in the case of adding storage nodes to expand storage space, the metadata node can be configured to: record the information of the newly added storage node, and in response to a request from a third computing node among the multiple computing nodes to store a data stream, allocate a storage location including a storage location located in the newly added storage node to the third computing node. In one embodiment, the control node can provide the metadata node with the information of the newly added storage node, that is, the control node can be configured to send the information of the newly added storage node to the metadata node based on the newly added storage node. In another embodiment, the newly added storage node can also send its own information to the metadata node.
[0086] Figure 6 A schematic diagram is shown of how to quickly perform reading and writing after a storage node is added according to an embodiment of the present disclosure.
[0087] Assuming that a new storage node 623 is added, compared to the traditional method of moving part of the data in the existing storage nodes 621 and 622 to the new storage node 623, according to the embodiment of the present disclosure, the data may not be moved, but the metadata node 630 responds to a request from a computing node 613 among the multiple computing nodes 611, 612, 613, and 614 to store a data stream, and allocates a storage location including a storage location in the newly added storage node 623 to the computing node 613, so as to store the data blocks 6231, 6232, and 6222 of the data stream. Figure 6As shown, data blocks 6231 and 6232 are located in the newly added storage node 623, while data block 6222 is located in the original storage node 622. Of course, all data blocks can also be stored in the newly added storage node 623. The metadata node 630 records the corresponding relationship 638 between each data block of the data stream and the storage location storing each data block of the data stream.
[0088] In addition, the metadata node 630 may also record multiple references 635 and 636 for the data stream, which are used by the computing nodes 613 and 614 to obtain the corresponding relationship 638 for the data stream to access the data stream.
[0089] Thus, according to the storage and reading operations of the data stream described previously, the storage operation of the data stream can be immediately performed in the newly added storage node, and the reading operation can be performed subsequently. Compared with the problem that the existing database architecture cannot immediately provide the corresponding write service when the storage node is expanded, according to the embodiments of the present disclosure, what is exposed to the computing node at the storage layer is an abstraction of the data block level, and a data block can be allocated to any storage location in any storage node for storage, so when a new storage node is added, this data block-level write service can be immediately provided.
[0090] Scenario 2 - Zero-copy dynamic load balancing
[0091] As data is continuously written, the distribution of data streams contained in certain databases (such as tables) in computing nodes will change, and the frequency of requests for different data will also change, resulting in some computing nodes possibly needing to process more data, thereby causing performance bottlenecks. According to the embodiments of the present disclosure, a dynamic load balancing mechanism can be executed without copying data blocks, thereby avoiding performance bottlenecks caused by data distribution.
[0092] Specifically, the control node is configured to: schedule the second data stream processed by the fourth computing node among the multiple computing nodes to be processed by the fifth computing node among the multiple computing nodes according to the load distribution of the multiple computing nodes, and control the metadata node to establish a reference to the second data stream for the fifth computing node during the scheduling, and delete the reference to the second data stream for the fourth computing node after the scheduling is completed.
[0093] For example, the control node can obtain the amount of data referenced by each computing node and the load it brings, which is recorded in real time by each computing node, so as to know the load distribution of multiple computing nodes. If the control node finds that this load distribution is not optimal, or there is an unbalanced load distribution, such as the load at some computing nodes is significantly higher than the load at other computing nodes, it is necessary to schedule some data flows (such as the second data flow) on those computing nodes with high loads (such as the fourth computing node) to other computing nodes for processing (such as the fifth computing node).
[0094] After the reference to the second data stream is established for the fifth computing node and the scheduling is completed, the fifth computing node can take over the second data stream without copying or moving the second data stream, and can delete the reference to the second data stream for the fourth computing node. In this way, relying on the efficient data stream reference mechanism provided by the metadata node, on the basis of zero copy of the data stream, the data stream can be quickly taken over by other computing nodes, thereby eliminating the performance bottleneck of the system.
[0095] In addition, during the scheduling process, since the original reference of the fourth computing node to the second data stream is still valid, the fourth computing node can still access the second data stream, so the scheduling process will not affect the processing of the fourth computing node's read request for the second data stream. After the scheduling is completed, the user's read request for the second data stream will be automatically routed to the new fifth computing node, and the old reference of the fourth computing node will be deleted. At this time, the fifth computing node completely takes over the second data stream.
[0096] Scenario 3: Compute node groups flexibly share and exclusively use database resources
[0097] Database resource sharing can reduce hardware and software costs, and exclusive use of resources can ensure service stability and provide higher security. According to the embodiments of the present disclosure, two databases (data stream sets) can be assigned to the same group of computing nodes so that they share the same set of computing resources (i.e., resource sharing). For example, references to the two databases can be read by the same computing node; the two databases can also be assigned to two different groups of computing nodes to physically ensure their isolation (i.e., exclusive use of resources). Here, a data entry in a database can be streamed into a data stream, so a database can be represented as a data stream set, and a data stream set can include one or more data streams.
[0098] According to the embodiments of the present disclosure, a database can be migrated from one group of computing nodes to another group of computing nodes with zero copy of data blocks, so that a quick switch can be made between shared resources and exclusive resources of the database.
[0099] Specifically, the control node can be configured to: control the metadata node to copy the reference to the data stream in the first data stream set for each computing node of the first computing node group used to process the first data stream set to create a reference to the data stream in the first data stream set for each computing node in the original second computing node group; and control the first data stream set to go offline from the first computing node group. The metadata node can be configured to: delete the reference to the data stream in the first data stream set for each computing node of the first computing node group. In this way, the first data stream set can be migrated from the first computing node group to the second computing node group with zero copy of data blocks, and the second computing node group accesses the first data stream set and provides services to the outside.
[0100] Scenario 4 - Database Fast Zero Copy Multi-copy Architecture
[0101] When the database still cannot meet user requests through horizontal expansion, vertical expansion is required. Traditional databases need to physically copy the data before expansion. This approach not only wastes storage space, but is also relatively slow. According to an embodiment of the present disclosure, the service capacity of the database can be increased by having multiple computing nodes reference the same data and provide services. In other words, according to an embodiment of the present disclosure, the purpose of increasing the service capacity of the processing database can be achieved by adding computing nodes without copying the data stream.
[0102] Figure 7 A schematic diagram of a database fast zero-copy multi-copy architecture according to an embodiment of the present disclosure is shown.
[0103] like Figure 7 As shown, first, the control node 740 can apply for cloud resources to select or construct a new computing node group 712, for example, the newly added second computing node group in the aforementioned embodiment. Specifically, the control node 740 can be configured to: control the metadata node 730 to copy the reference of each computing node of the first computing node group 711 for processing the first data stream set (first database) to the data stream in the first data stream set to create a reference to the data stream in the first data stream set for each computing node in the newly added second computing node group 712.
[0104] At the same time, in order to enable the newly added second computing node group 712 to obtain complete information of the first computing node group 711, the control node is also configured to: control each computing node in the first computing node group 711 to synchronize the data in the memory to the memory of the computing node in the second computing node group 712.
[0105] For example, the memory of the first computing node group 711 may be loaded with data to be stored in the storage node. In order to completely synchronize the current process of the first computing node group 711, the second computing node group 712 can obtain the data in the memory of the first computing node group 711 and store it in its own memory, so as to perform database operations based on the data in the memory. After the memory of the second computing node group 712 obtains the data in the memory of the first computing node group 711, it can directly read the data from its own memory, thereby improving the data reading speed. For example, the synchronization method can be: the control node 740 controls the second computing node group 712 to obtain the current memory data from each computing node of the first computing node group 711 by sending a network request, and synchronously loads it into the memory of the second computing node group 712.
[0106] In this way, since the metadata node 730 simultaneously records the reference of the first computing node group 711 to the first data stream set and the reference of the second computing node group 712 to the first data stream set, and the real-time data in the memory is also synchronized, the first computing node group 711 and the second computing node group 712 can both process the first data stream set, thereby realizing fast zero-copy multiple copies of the database.
[0107] Scenario 5 - Zero-copy fast cloning of a readable and writable database
[0108] When using a database, users hope to be able to quickly conduct multiple different experiments in a piece of data, and hope that the experiments will not interfere with each other. According to the embodiments of the present disclosure, a readable and writable database can be quickly cloned with zero copy, and provided to different business parties for simultaneous experiments. Users can continue to initiate new read and write requests on the new cloned database, and can also continue to clone new databases. This user experience greatly improves the efficiency of business iteration and reduces the cost of use.
[0109] According to the embodiments of the present disclosure, in order to quickly clone a readable and writable database with zero copy, it is only necessary to copy a new reference to all data streams in the database, and the new reference is presented as a new data block from the perspective of the computing node. When the new data block needs to be loaded to a computing node, it is only necessary to assign the corresponding reference to the computing node.
[0110] Scenario 6 - Smooth upgrade based on fast database cloning
[0111] Smooth database upgrades can ensure the normal operation of the system, reduce downtime and the risk of data loss, improve system availability and user experience, and also bring performance and functional improvements. However, in the existing technology, there is no particularly good solution for system upgrades that require collaborative services for the status of each node, which can both guarantee a high service level agreement (SLA) and reduce data storage and management costs.
[0112] Figure 8 A schematic diagram of a smooth upgrade process based on fast database cloning according to an embodiment of the present disclosure is shown.
[0113] like Figure 8 As shown, the control node 860 may receive an instruction from the client 850 and determine to upgrade from the first computing node group 811 to the second computing node group 812. The second computing node group 812 may be a group of computing nodes newly applied by the control node 860 from the cloud service.
[0114] Similar to scenario 4, the control node 860 can be configured to: control the metadata node to copy the reference to the data stream in the first data stream set of each computing node in the first computing node group 811 for processing the first data stream set to create a reference to the data stream in the first data stream set for each computing node in the newly added second computing node group 812; control each computing node in the first computing node group 811 to synchronize the data in the memory to the memory of the computing node in the second computing node group 812.
[0115] In order to improve user experience and smooth database upgrade, the control node 860 can also be configured to: when the difference between the data in the memory of the first computing node group 811 and the data in the memory of the second computing node group 812 is less than a first predetermined threshold or the synchronization delay between the first computing node group and the second computing node group is lower than a second predetermined threshold, control the first computing node group 811 to stop writing data to its memory, control the second computing node group 812 to go online after the synchronization is completed, and control the first computing node group 811 to go offline, and control the metadata node to delete the reference to the data stream in the first data stream set for each computing node of the first computing node group 811 after the synchronization is completed. In an embodiment of the present disclosure, the online of a computing node indicates that the computing node can perform read and write operations on the corresponding database.
[0116] The first predetermined threshold or the second predetermined threshold determines the time during which write operations are not allowed. The larger the threshold, the larger the amount of data that needs to be synchronized after writing stops, and the longer the synchronization time; the smaller the threshold, the smaller the amount of data that needs to be synchronized after writing stops, and the shorter the synchronization time. After the synchronization is completed, the second computing node group goes online, and subsequent data can be written through the second computing node group, for example, first written to the memory of the second node group and then stored in the storage node, or directly stored in the storage node. Therefore, from the user's perspective, write operations cannot be performed only from the time the first computing node group stops writing to the time the second computing node group completes synchronization and goes online. Therefore, by setting a suitable threshold, users can be unable to perform write operations for only a very short period of time, so that users cannot even perceive the stop process, thereby improving user experience.
[0117] For example, when the difference between the data in the memory of the first computing node group 811 and the data in the memory of the second computing node group 812 is small, since the amount of data that differs in the memory of the two computing node groups is small, the write process is stopped at this time, so that the first computing node group can quickly synchronize the data with less difference in memory to the memory of the second computing node group. As a result, the second computing node group quickly goes online after the synchronization is completed, and then performs subsequent data write operations. By setting a suitable first threshold value, the process time is very short, and the user may not even perceive the process of stopping writing, thereby bringing a better user experience to the user.
[0118] For another example, if the synchronization delay between the first computing node group 811 and the second computing node group 812 is low, the data in the memory of the first computing node group 811 and the data in the memory of the second computing node group 812 are small in difference because they can be quickly synchronized, and it is appropriate to stop the write operation of the first computing node group 811 at this time. In addition, due to the small synchronization delay, the speed at which the difference data in the memory is further synchronized is also fast, so the subsequent further synchronization time is shorter. Therefore, at this time, it is appropriate to stop the write operation of the first computing node group 811 and complete the synchronization of the first computing node group 811 and the second computing node group 812, so as to bring a better user experience to the user.
[0119] The first threshold and the second threshold can be set based on experience or experimentation based on system requirements. For example, the data difference or synchronization delay between the two computing node groups is determined as the first threshold or the second threshold based on the maximum time length (e.g., 1 second) of the expected stop write operation and with reference to historical experience values.
[0120] In addition, the difference between the data in the memory of the first computing node group and the data in the memory of the second computing node group, or the synchronization delay between the first computing node group and the second computing node group can be determined by various technical means. For example, each computing node of the computing node group can monitor the data written into the memory and report the corresponding data to the aggregation node (such as a control node or a computing node), so that the aggregation node can determine the data difference in the memory of each computing node. For another example, the synchronization delay can be determined by the difference between the moment when a computing node of the first computing node group sends data and the moment when a computing node of the second computing node group receives the data. In an embodiment of the present disclosure, the synchronization delay represents the time difference from when data is sent from one node to when the data is completely received by another node, including the decoding and processing time when the data is received.
[0121] It should be noted that in the embodiments of the present disclosure, the components and operations of the "computing node group" represent the components and operations of the computing nodes in the computing node group. For example, the memory of the "first / second computing node group" represents the memory of the computing nodes in the "first / second computing node group", which can be the memory of all computing nodes or the memory of representative computing nodes. For another example, the synchronization delay between the first computing node group and the second computing node group can be the average of the synchronization delays between all corresponding computing nodes in the two computing node groups, or it can be the synchronization delay between one or more representative computing nodes in the two computing node groups.
[0122] Through the above process, the database service is smoothly upgraded from the first computing node group 811 to the second computing node group without actually copying the first data stream set (data block). In the above entire upgrade process, the database reading service can be provided to the user all the time, and the write processing is only briefly stopped, about a second level, so the user hardly feels the lag, which brings a good experience to the user.
[0123] Scenario 7 - Rapid backup and recovery of database
[0124] Database backup is an important means to protect data security, provide data recovery and disaster recovery capabilities, and is crucial to ensuring business continuity and data reliability.
[0125] Traditional database backup requires physical copying of data, which greatly wastes storage space. At the same time, the backed-up data also needs to be copied again during the recovery phase, which takes a long time to recover.
[0126] According to the embodiments of the present disclosure, there is no need to actually copy the data. Backing up can be done by simply adding a reference to the data, which greatly saves storage space. In addition, during the database recovery process, the same reference technology can be used to associate the computing nodes with the references and quickly load the backup into the computing node group, thereby greatly accelerating the database recovery process.
[0127] According to an embodiment of the present disclosure, a metadata node for a distributed database architecture system is also provided, the distributed database architecture system comprising multiple computing nodes, multiple storage nodes and the metadata node, wherein the metadata node is configured to: record the correspondence between each data stream and one or more data block resources storing each data stream, the one or more data block resources are located in one or more of the multiple storage nodes, and each of the one or more data block resources stores one of the one or more data blocks divided by each data stream; record more than two references to a first data stream to be read by more than two computing nodes, each reference is used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream.
[0128] The description of metadata nodes in the above description of the distributed database architecture system is also applicable here and will not be repeated here.
[0129] Fig. 9 A flow chart of a method 900 for a distributed database architecture system according to an embodiment of the present disclosure is shown. The distributed database architecture system includes a plurality of computing nodes, a plurality of storage nodes, and a metadata node. Fig. 9 As shown, the method 900 includes steps 910 to 930.
[0130] In step 910, the metadata node records the correspondence between each data block in the one or more data blocks into which each data stream is divided and the storage location used to store the each data block, and the storage location of the one or more data blocks for each data stream is located in one or more of the multiple storage nodes; and more than two references are recorded for a first data stream to be read by more than two computing nodes, and each reference is used by one of the more than two computing nodes to obtain the correspondence for the first data stream to read the first data stream.
[0131] In step 920, the plurality of computing nodes read each data stream from the plurality of storage nodes or store each data stream to the plurality of storage nodes according to the corresponding relationship for each data stream.
[0132] In step 930, the plurality of storage nodes store the one or more data blocks of each data stream into storage locations corresponding to the one or more data blocks.
[0133] In some embodiments, the method 900 also includes: in response to a first computing node among the more than two computing nodes needing to read the first data stream, the metadata node generates a reference to the first data stream for the first computing node; in response to the first computing node no longer needing to read the first data stream, the metadata node deletes the reference to the first data stream for the first computing node; and in response to all references to the first data stream being deleted, instructing the multiple storage nodes to delete the first data stream.
[0134] In some embodiments, the method 900 also includes: the metadata node allocates one or more first storage locations located in one or more first storage nodes among the multiple storage nodes to the one or more first data blocks into which the first data stream is divided, in response to a storage request from a second computing node among the multiple computing nodes to store the first data stream, returns information about the one or more first storage locations for storing the one or more first data blocks to the second computing node, and records the correspondence between each first data block in the one or more first data blocks and the first storage location for storing the each first data block; the second computing node divides the first data stream into the one or more first data blocks, and sends the corresponding first data block in the one or more first data blocks to the first storage node corresponding to each first storage location according to the information of the one or more first storage locations; and the first storage node stores the first data block received from the second computing node in the corresponding first storage location.
[0135] In some embodiments, the method 900 also includes: the metadata node records information of the newly added storage node, and in response to a request from a third computing node among the multiple computing nodes to store a data stream, allocates a storage location including a storage location located in the newly added storage node to the third computing node.
[0136] In some embodiments, the distributed database architecture system further includes a control node, wherein the method 900 further includes: the control node sending information of the newly added storage node to the metadata node based on the newly added storage node.
[0137] In some embodiments, the distributed database architecture system also includes a control node, wherein the method 900 also includes: the control node: scheduling the second data stream processed by a fourth computing node among the multiple computing nodes to a fifth computing node among the multiple computing nodes for processing according to the load distribution of the multiple computing nodes, and controlling the metadata node to establish a reference to the second data stream for the fifth computing node during the scheduling, and deleting the reference to the second data stream for the fourth computing node after the scheduling is completed.
[0138] In some embodiments, the distributed database architecture system also includes a control node, wherein the method 900 also includes: the control node controls the metadata node to copy the reference to the data stream in the first data stream set for each computing node in the first computing node group used to process the first data stream set to create a reference to the data stream in the first data stream set for each computing node in the newly added or original second computing node group.
[0139] In some embodiments, the method 900 also includes: controlling, by the control node, the first data stream set to be offline from the first computing node group; and deleting, by the metadata node, references to data streams in the first data stream set for each computing node in the first computing node group.
[0140] In some embodiments, the method 900 further includes: the control node controls each computing node in the first computing node group to synchronize the data in the memory to the memory of the computing node in the second computing node group.
[0141] In some embodiments, the method 900 also includes: by the control point: when the difference between the data in the memory of the first computing node group and the data in the memory of the second computing node group is less than a first predetermined threshold or the synchronization delay between the first computing node group and the second computing node group is lower than a second predetermined threshold, controlling the first computing node group to stop writing data to its memory; after the synchronization is completed, controlling the second computing node group to go online, and controlling the first computing node group to go offline; and after the synchronization is completed, controlling the metadata node to delete the reference to the data stream in the first data stream set for each computing node of the first computing node group.
[0142] In some embodiments, the method 900 further includes: storing, by the metadata node, multiple copies of each reference.
[0143] In some embodiments, the plurality of computing nodes are deployed in one or more cloud services; and / or the plurality of storage nodes are deployed in one or more cloud services.
[0144] The above description on the distributed database architecture system is also applicable to the method for the distributed database architecture system.
[0145] Fig.10 A flowchart of a method 1000 for a metadata node in a distributed database architecture system is shown. The distributed database architecture system includes a plurality of computing nodes, a plurality of storage nodes, and a metadata node. Fig.10 As shown, the method 1000 includes steps 1010 to 1020.
[0146] In step 1010, the metadata node records the correspondence between each data block in the one or more data blocks into which each data stream is divided and the storage location used to store the each data block, and the storage location of the one or more data blocks for each data stream is located in one or more of the multiple storage nodes.
[0147] In step 1020, the metadata node records more than two references to a first data stream to be read by more than two computing nodes, and each reference is used by one of the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream.
[0148] The above descriptions of the distributed database architecture system, metadata node, and method for the distributed database architecture system are also applicable to the method for the metadata node.
[0149] In summary, according to the various embodiments of the present disclosure, the unbinding and flexible scheduling of computing nodes and storage nodes are realized, the stored data stream can be quickly read, and different computing nodes can independently access the same data stream, thereby realizing flexible deployment and rapid upgrade of the database. According to the various embodiments of the present disclosure, it is also possible to solve the rapid expansion of storage nodes, dynamic load balancing of computing nodes without copying data, resource sharing and exclusive use of databases, rapid zero-copy multiple copies of database granularity, rapid cloning of readable and writable databases without copying data, rapid and smooth upgrade of database clones, rapid backup and recovery of database granularity, etc.
[0150] It should be noted that the above-mentioned specific embodiments are only examples and not limitations, and those skilled in the art can merge and combine some steps and devices from the various embodiments described separately according to the concept of the present disclosure to achieve the effects of the present disclosure. Such merged and combined embodiments are also included in the present disclosure, and such merges and combinations are not described one by one here. In addition, the "first", "second", "third", and "fourth" in the present disclosure are only for marking one or some elements, and do not indicate their importance or particularity, nor do they indicate their differences. For example, the first computing node and the second computing node can be different computing nodes or the same computing nodes.
[0151] The advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and facilitation of understanding, not limitation, and the above details do not limit the present disclosure to being implemented by adopting the above specific details.
[0152] The block diagrams of devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems may be connected, arranged, or configured in any manner.
[0153] The step flow charts and the above method descriptions in this disclosure are intended only as illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the order given. As will be appreciated by those skilled in the art, the order of the steps in the above embodiments may be performed in any order. Words such as "thereafter," "then," "next," and the like are not intended to limit the order of the steps; these words are merely used to guide the reader through the descriptions of these methods.
[0154] Each operation of the method described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including but not limited to hardware circuits, application specific integrated circuits (ASICs) or processors.
[0155] The present disclosure may also include computer program products, wherein the computer program products can perform the methods, steps, and operations presented herein. For example, such computer program products can be computer software packages, computer code instructions, computer-readable tangible media having computer instructions tangibly stored (and / or encoded) thereon, which instructions can be executed by a processor to perform the operations described herein.
Claims
1. A distributed database architecture system, comprising a plurality of computing nodes, a plurality of storage nodes and a metadata node, wherein The metadata node is configured as: Recording a correspondence between each data block in the one or more data blocks into which each data stream is divided and a storage location for storing the each data block, wherein the storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes, and Recording more than two references to a first data stream to be read by more than two computing nodes, each reference being used by one computing node among the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream; The plurality of computing nodes are configured as: Reading each data stream from the plurality of storage nodes or storing each data stream to the plurality of storage nodes according to the corresponding relationship for each data stream; and The plurality of storage nodes are configured as follows: The one or more data blocks of each data stream are stored in storage locations for the one or more data blocks, respectively.
2. The distributed database architecture system according to claim 1, wherein: The metadata node is also configured to: In response to a first computing node among the two or more computing nodes needing to read the first data stream, generating a reference to the first data stream for the first computing node; In response to the first computing node no longer needing to read the first data stream, deleting a reference to the first data stream for the first computing node; as well as In response to all references to the first data stream being deleted, instructing the plurality of storage nodes to delete the first data stream.
3. The distributed database architecture system according to claim 1, wherein The metadata node is also configured to: In response to a storage request from a second computing node among the plurality of computing nodes to store the first data stream, allocating one or more first storage locations in one or more first storage nodes among the plurality of storage nodes to the one or more first data blocks into which the first data stream is divided, returning information of the one or more first storage locations for storing the one or more first data blocks to the second computing node, and Recording a correspondence between each first data block in the one or more first data blocks and a first storage location for storing each first data block; The second computing node is configured as: dividing the first data stream into the one or more first data blocks, and Sending a corresponding first data block among the one or more first data blocks to a first storage node corresponding to each first storage location according to information of the one or more first storage locations; as well as The first storage node is configured as: The first data block received from the second computing node is stored in a corresponding first storage location.
4. The distributed database architecture system according to claim 1, wherein The metadata node is also configured to record information of a newly added storage node, and in response to a request from a third computing node among the multiple computing nodes to store a data stream, allocate a storage location including a storage location located in the newly added storage node to the third computing node.
5. The distributed database architecture system according to claim 4 further comprises a control node, wherein The control node is configured to send information of the newly added storage node to the metadata node according to the newly added storage node.
6. The distributed database architecture system according to claim 1 further comprises a control node, wherein: The control node is configured to: According to the load distribution of the plurality of computing nodes, scheduling the second data flow processed by a fourth computing node among the plurality of computing nodes to be processed by a fifth computing node among the plurality of computing nodes, and Control the metadata node to establish a reference to the second data stream for the fifth computing node during the scheduling, and delete the reference to the second data stream for the fourth computing node after the scheduling is completed.
7. The distributed database architecture system according to claim 1 further comprises a control node, wherein The control node is configured to: Control the metadata node to copy the reference to the data stream in the first data stream set for each computing node in the first computing node group for processing the first data stream set to create a reference to the data stream in the first data stream set for each computing node in a newly added or existing third computing node group.
8. The distributed database architecture system according to claim 7, wherein The control node is further configured to: control the first data flow set to be offline from the first computing node group; and The metadata node is also configured to delete references to data flows in the first set of data flows for each compute node of the first group of compute nodes.
9. The distributed database architecture system according to claim 7, wherein The control node is further configured to: control each computing node in the first computing node group to synchronize the data in the memory to the memory of the computing node in the second computing node group.
10. The distributed database architecture system according to claim 9, wherein The control node is further configured to: When the difference between the data in the memory of the first computing node group and the data in the memory of the second computing node group is less than a first predetermined threshold or the synchronization delay between the first computing node group and the second computing node group is lower than a second predetermined threshold, control the first computing node group to stop writing data to its memory, After the synchronization is completed, controlling the second computing node group to go online, and controlling the first computing node group to go offline, and After the synchronization is completed, the metadata node is controlled to delete references to data streams in the first data stream set for each computing node of the first computing node group.
11. The distributed database architecture system according to claim 1, wherein The metadata node is also configured to store multiple copies of each reference.
12. The distributed database architecture system according to claim 1, wherein The plurality of computing nodes are deployed in one or more cloud services; and / or The multiple storage nodes are deployed in one or more cloud services.
13. A metadata node for a distributed database architecture system, the distributed database architecture system comprising a plurality of computing nodes, a plurality of storage nodes and the metadata node, wherein the metadata node is configured as: Recording a correspondence between each data block in the one or more data blocks into which each data stream is divided and a storage location for storing the each data block, wherein the storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes; and More than two references are recorded for a first data stream to be read by more than two computing nodes, and each reference is used by one computing node among the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream.
14. A method for a distributed database architecture system, the distributed database architecture system comprising a plurality of computing nodes, a plurality of storage nodes and a metadata node, wherein: The method comprises: From the metadata node: Recording a correspondence between each data block in the one or more data blocks into which each data stream is divided and a storage location for storing the each data block, wherein the storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes, and Recording more than two references to a first data stream to be read by more than two computing nodes, each reference being used by one computing node among the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream; The plurality of computing nodes: Reading each data stream from the plurality of storage nodes or storing each data stream to the plurality of storage nodes according to the corresponding relationship for each data stream; and The plurality of storage nodes: The one or more data blocks of each data stream are stored in storage locations for the one or more data blocks, respectively.
15. A method for a metadata node of a distributed database architecture system, the distributed database architecture system comprising a plurality of computing nodes, a plurality of storage nodes and the metadata node, wherein the method comprises: Recording a correspondence between each data block in the one or more data blocks into which each data stream is divided and a storage location for storing each data block, wherein the storage location of the one or more data blocks for each data stream is located in one or more of the plurality of storage nodes; as well as More than two references are recorded for a first data stream to be read by more than two computing nodes, and each reference is used by one computing node among the more than two computing nodes to obtain the corresponding relationship for the first data stream to read the first data stream.