Database cluster, data processing method, data processing system and related device

By introducing storage engine instances to the database cluster to share metadata, the problem of high resource consumption when synchronizing metadata is solved, and higher performance and better storage resource utilization are achieved.

WO2025118658A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109959
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-08-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Database clusters consume high resources when synchronizing metadata, which affects performance.

Method used

By introducing storage engine instances into the database cluster, metadata in the storage device is shared, so that the service instance does not need to synchronously update metadata in the local storage area.

Benefits of technology

It avoids resource consumption caused by synchronizing metadata in the database cluster, improves the performance of the database cluster, and optimizes the utilization of storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109959_12062025_PF_FP_ABST
    Figure CN2024109959_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of databases. Provided are a database cluster, a data processing method, a data processing system and a related device. The database cluster comprises a first service instance, a second service instance and a storage engine instance, the storage engine instance being connected to the first service instance, the second service instance and a storage apparatus, and metadata in the storage apparatus being shared by the first service instance and the second service instance. The first service instance is used for sending an operation command to the storage engine instance on the basis of a database statement; and the storage engine instance is used for updating first data and second metadata which are persistently stored in the storage apparatus on the basis of the operation command. Therefore, a plurality of service instances can share the same metadata in a storage apparatus by means of a storage engine instance, without the need of independently and persistently storing the metadata in a local storage area, so that after the metadata in the storage apparatus is updated, there is no need to execute a metadata synchronization process, thereby improving the performance of a database cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Database cluster, data processing method, data processing system and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 8, 2023, with application number 202311691625.2 and application name “Database Cluster, Data Processing Method, Data Processing System and Related Equipment”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of database technology, and in particular to a database cluster, a data processing method, a data processing system and related equipment. Background Art

[0003] With the development of information technology, database systems, such as MySQL, are widely used in finance, communications, medical care, logistics, e-commerce and other fields, and are used to add, delete, modify, and query business data in various fields.

[0004] Currently, database systems are facing increasing demands for concurrent read and write operations and availability. Consequently, multiple database systems are often integrated in real-world applications to form a database cluster. Each database system stores a persistent copy of metadata locally. This allows multiple database systems to leverage their respective metadata to access persistently stored data in the same storage area (such as a cloud storage area), improving the concurrent read and write capabilities and availability of the database cluster.

[0005] However, when a database system modifies the data in the storage area, the database system not only needs to update the metadata stored in itself, but also needs to synchronize the updated metadata to other database systems to ensure the consistency of the metadata stored locally by multiple database systems. This will result in high resource consumption of the database cluster and affect the performance of the database cluster.

[0006] Summary of the Invention

[0007] The present application provides a database cluster to improve the performance of the database cluster. In addition, the present application also provides a data processing method, a data processing system, a device cluster, a computer-readable storage medium, and a computer program product.

[0008] In a first aspect, the present application provides a database cluster comprising multiple service instances and at least one storage engine instance. For example, the service instances in the database cluster may be service instances based on a MySQL database, and the storage engine instances may be storage engine instances based on a MySQL database. The at least one storage engine instance is connected to the multiple service instances and a storage device, and the storage device is configured to persistently store first data and first metadata corresponding to the first data. The multiple service instances include a first service instance and a second service instance, and the first storage engine instance in the at least one storage engine instance is connected to the first service instance and the second service instance. The first metadata is shared by the first and second service instances, meaning that both the first and second service instances can access the first metadata in the storage device. The first service instance is configured to obtain a first database statement and, based on the first database statement, send a first operation command to the first storage engine instance. The first database statement may be, for example, a MySQL statement, and is configured to instruct an update of the first data. The first storage engine instance is configured to update the first data in the storage device based on the first operation command sent by the first service instance, thereby obtaining second data, and then update the first metadata to the second metadata corresponding to the second data.

[0009] Because the first metadata is stored in the storage device and the storage engine instance can access the first metadata in the storage device, the first service instance and the second service instance connected to the first storage engine instance can share the same metadata in the storage device through the first storage engine instance. This allows the first service instance to update the first metadata corresponding to the first data in the storage device to the second metadata corresponding to the second data through the first storage engine instance without having to perform a metadata synchronization process, that is, it is not necessary to synchronize the second metadata to the local storage area (for persistent storage) of each service instance. This can avoid the resource consumption caused by synchronizing metadata in the database cluster, and the database cluster can use more resources to process business, thereby improving the performance of the database cluster. In addition, the first service instance and the second service instance can share the same metadata in the storage device through the first storage engine instance, which eliminates the need for the first service instance and the second service instance to separately and persistently store a copy of the metadata in the local storage area, thereby avoiding the storage resource consumption caused by persistently storing multiple copies of metadata in the database cluster. In this way, the database cluster can use more storage resources to process business, thereby further improving the performance of the database cluster.

[0010] In one possible embodiment, at least one storage engine instance in the database cluster further includes a second storage engine instance, and the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device. The third service instance is used to obtain a second database statement and, based on the second database statement, send a second operation command to the second storage engine instance, wherein the second database statement is used to update the second data. The second storage engine instance is used to further update the second data according to the second operation command to obtain third data, and update the second metadata to third metadata corresponding to the third data. In this way, multiple service instances can update the data and metadata in the storage device through their respective connected storage engine instances, and the updated metadata does not need to be synchronized between multiple service instances, thereby avoiding resource consumption caused by synchronizing metadata in the database cluster and improving the performance of the database cluster.

[0011] In one possible embodiment, the storage device includes multiple sub-storage devices, taking the first sub-storage device and the second sub-storage device as an example, the first sub-storage device is used to store first data and first metadata, and the second sub-storage device and the first sub-storage device are used to store different data and different metadata, such as data of different shards and metadata corresponding to the data of the shards. At least one storage engine instance in the database cluster also includes a second storage engine instance, which is connected to the first service instance, the second service instance and the second sub-storage device. The first storage engine instance is specifically connected to the first sub-storage device. At this time, the first storage engine instance is used to access the first sub-storage device, and the second storage engine instance is used to access the second sub-storage device. In this way, the storage device can use a distributed storage structure to store data and metadata, so that the service instance can use different storage engine instances to update the data and metadata in different sub-storage devices, and there is no need to synchronize metadata in the database cluster, thereby improving the performance of the database cluster.

[0012] In one possible implementation, a storage device includes a metadata storage device and multiple data storage devices. For example, a first data storage device and a second data storage device are included. The first data storage device is used to store first data, and the first and second data storage devices are used to store different data, such as data from different shards. The metadata storage device is used to store metadata corresponding to all data, including metadata corresponding to the data stored by the first data storage device and metadata corresponding to the data stored by the second data storage device. At least one storage engine instance in the database cluster also includes a second storage engine instance, which is connected to the first service instance, the second service instance, and the second data storage device. The first storage engine instance is specifically connected to the first data storage device, and both the first and second storage engine instances are connected to the metadata storage device. In this case, the first storage engine instance is used to access the first data storage device and the metadata storage device, while the second storage engine instance is used to access the second data storage device and the metadata storage device. In this manner, the storage device can use a distributed storage structure to store data, while metadata can be stored using a centralized storage method. This allows service instances to use different storage engine instances to update data and metadata in different data storage devices without synchronizing metadata across the database cluster, thereby improving the performance of the database cluster.

[0013] In one possible implementation, the first service instance is further used to set a lock for the first metadata before sending the first operation command to the first storage engine instance, so as to use the set lock to limit the reading and writing of the first data corresponding to the first metadata, and release the lock set for the first metadata after the first metadata is updated to the second metadata. The first storage engine instance is also used to instruct the second service instance connected to the first storage engine instance to set a lock for the first metadata, and after updating the first metadata to the second metadata, instruct the second service instance to release the lock set for the first metadata. In this way, when the first service instance updates the data and metadata in the storage device through the first storage engine instance, the first service instance and the second service instance can prohibit the first service instance and the second service instance from updating (and reading) the data described by the metadata by setting a lock for the first metadata, thereby ensuring that multiple service instances will not modify the same data and the same metadata concurrently.

[0014] In one possible implementation, at least one storage engine instance in a database cluster also includes a second storage engine instance, and the multiple service instances also include a third service instance. Furthermore, the second storage engine instance is connected to the third service instance and a storage device. The first storage engine instance is further configured to, before updating the first metadata to the second metadata, instruct the third service instance, via the second storage engine instance, to set a lock for the first metadata. Furthermore, after updating the first metadata to the second metadata, instruct the third service instance, via the second storage engine instance, to release the lock set for the first metadata. Thus, when a database cluster contains multiple storage engine instances, the first service instance can, through these multiple storage engine instances, notify the remaining service instances to set a lock for the first metadata, thereby ensuring that multiple service instances do not concurrently modify the same data and metadata.

[0015] In one possible implementation, before the first metadata is updated to the second metadata, the first metadata is cached in the first service instance, the second service instance, and the third service instance, so as to improve the parsing efficiency of database statements by utilizing the cached first metadata. The first service instance is also used to invalidate the first metadata cached by the first service instance after the first metadata is updated to the second metadata; and the first storage engine instance is also used to instruct the second service instance to invalidate the first metadata cached by the second service instance, and to instruct the third service instance to invalidate the first metadata cached by the third service instance through the second storage engine instance. In this way, the first service instance notifies the remaining service instances through multiple storage engine instances to invalidate the cached first metadata, thereby preventing each service instance from subsequently executing an erroneous data processing process based on the expired first metadata.

[0016] In one possible implementation, the second service instance is configured to cache second metadata from a storage device after invalidating the cached first metadata. By re-caching the metadata from the storage device, the second service instance can parse subsequently acquired database statements using the latest metadata, improving the efficiency and accuracy of database statement parsing. Furthermore, multiple service instances can re-cache metadata from the storage device, thereby ensuring consistent cached metadata across multiple service instances.

[0017] In one possible implementation, the first service instance is further configured to send a transaction to the first storage engine instance after the first metadata is updated to the second metadata. The transaction includes the database statement obtained by the first service instance. The first storage engine instance is further configured to commit the transaction. In this way, the database cluster can modify the data and metadata in the storage device within a single transaction, ensuring the atomicity of the data and metadata modifications.

[0018] In one possible implementation, if the first storage engine instance fails before a transaction is committed, the second storage engine instance is used to roll back updates to the first data and first metadata and instruct the third service instance to release the lock set for the first metadata. In this way, by taking over the second storage engine instance, failure recovery can be achieved for the storage engine instance. Specifically, by rolling back updates to the first data and second metadata, the database cluster can be restored to its state before the failure of the first storage engine instance.

[0019] In one possible implementation, if the first storage engine instance fails after a transaction is committed, the second storage engine instance instructs the third service instance to release the lock set for the first metadata and invalidate the first metadata cached by the third service instance. This allows the second storage engine instance to take over and continue the unfinished operations of the first storage engine instance, completing the fault handling for the first storage engine instance in the database cluster and ensuring normal operation of the database cluster.

[0020] In a second aspect, the present application provides a data processing method, which is applied to a database cluster, wherein the database cluster includes multiple service instances and at least one storage engine instance, the at least one storage engine instance is connected to the multiple service instances and a storage device, the storage device is used to persistently store first data and first metadata corresponding to the first data, the first storage engine instance in the at least one storage engine instance is connected to the first service instance and the second service instance in the multiple service instances, and the first metadata is shared by the first service instance and the second service instance; when performing data processing in the database cluster, the first service instance obtains a first database statement, and the first database statement is used to indicate an update to the first data; then, the first service instance sends a first operation command to the first storage engine instance according to the first database statement; the first storage engine instance updates the first data according to the first operation command to obtain second data, and updates the first metadata to the second metadata corresponding to the second data.

[0021] In one possible implementation, at least one storage engine instance further includes a second storage engine instance, and the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; the method further includes: the third service instance obtains a second database statement, and the second database statement is used to instruct to update the second data; the third service instance sends a second operation command to the second storage engine instance according to the second database statement; the second storage engine instance updates the second data according to the second operation command to obtain third data; and the second storage engine instance updates the second metadata to third metadata corresponding to the third data.

[0022] In one possible embodiment, the storage device includes a first sub-storage device and a second sub-storage device, the first sub-storage device is used to store first data and first metadata, and the first sub-storage device and the second sub-storage device are used to store different data and different metadata; at least one storage engine instance also includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second sub-storage device, and the first storage engine instance is connected to the first sub-storage device; the first storage engine instance is used to access the first sub-storage device, and the second storage engine instance is used to access the second sub-storage device.

[0023] In one possible embodiment, the storage device includes a first data storage device, a second data storage device and a metadata storage device, the first data storage device is used to store first data, the first data storage device and the second data storage device are used to store different data, and the metadata storage device is used to store metadata corresponding to the data in the first data storage device and metadata corresponding to the data in the second data storage device; at least one storage engine instance also includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second data storage device, the first storage engine instance is connected to the first data storage device, and the first storage engine instance and the second storage engine instance are both connected to the metadata storage device; the first storage engine instance is used to access the first data storage device and the metadata storage device, and the second storage engine instance is used to access the second data storage device and the metadata storage device.

[0024] In one possible implementation, the method further includes: the first service instance sets a lock for the first metadata before sending the first operation command to the first storage engine instance; the first storage engine instance instructs the second service instance to set a lock for the first metadata; the first service instance releases the lock set for the first metadata after the first metadata is updated to the second metadata; and the first storage engine instance instructs the second service instance to release the lock set for the first metadata after updating the first metadata to the second metadata.

[0025] In one possible implementation, at least one storage engine instance further includes a second storage engine instance, and the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; the method further includes: before the first storage engine instance updates the first metadata to the second metadata, the first storage engine instance instructs the third service instance through the second storage engine instance to set a lock for the first metadata; after the first storage engine instance updates the first metadata to the second metadata, the first storage engine instance instructs the third service instance through the second storage engine instance to release the lock set for the first metadata.

[0026] In one possible implementation, before the first metadata is updated to the second metadata, the first metadata is cached in the first service instance, the second service instance, and the third service instance; the method further includes: after the first metadata is updated to the second metadata, the first service instance invalidates the first metadata cached by the first service instance; the first storage engine instance instructs the second service instance to invalidate the first metadata cached by the second service instance, and instructs the third service instance through the second storage engine instance to invalidate the first metadata cached by the third service instance.

[0027] In a possible implementation, the method further includes: after invalidating the cached first metadata, the second service instance caches the second metadata in the storage device.

[0028] In a possible implementation, the method further includes: after the first metadata is updated to the second metadata, the first service instance sends a transaction to the first storage engine instance, where the transaction includes a database statement; and the first storage engine instance commits the transaction.

[0029] In one possible implementation, the first storage engine instance fails before committing a transaction; the method further includes: the second storage engine instance rolls back the update operation on the first data and the first metadata; the second storage engine instance instructs the third service instance to release the lock set for the first metadata.

[0030] In one possible implementation, the first storage engine instance fails after a transaction is committed; the method further includes: the second storage engine instance instructs the third service instance to release the lock set for the first metadata; and the second storage engine instance invalidates the first metadata cached by the third service instance.

[0031] The data processing method provided in the second aspect corresponds to the database cluster provided in the first aspect. Therefore, the technical effects of the second aspect and any implementation method of the second aspect can be referred to the technical effects of the above-mentioned first aspect and the corresponding implementation method of the first aspect, and will not be repeated here.

[0032] In a third aspect, the present application provides a data processing system, which includes the database cluster described in the first aspect and any implementation of the first aspect, and the storage device described in the first aspect and any implementation of the first aspect.

[0033] In a fourth aspect, the present application provides a device cluster, which includes at least one computing node and at least one storage node. The at least one computing node is used to implement the database cluster described in the above-mentioned first aspect and any one of the implementation methods of the first aspect, and the at least one storage node is used to implement the storage device described in the above-mentioned first aspect and any one of the implementation methods of the first aspect.

[0034] In a fifth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computing device, the computing device executes the operating steps of the data processing method described in the second aspect or any implementation of the second aspect.

[0035] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device, enables the computing device to execute the operating steps of the data processing method described in the second aspect or any one of the implementations of the second aspect.

[0036] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a schematic diagram of the structure of an exemplary data processing system 10 provided by the present application;

[0038] FIG2 is a schematic structural diagram of another exemplary data processing system 10 provided by the present application;

[0039] FIG3 is a schematic structural diagram of another exemplary data processing system 10 provided by the present application;

[0040] FIG4 is a schematic structural diagram of another exemplary data processing system 10 provided in this application;

[0041] FIG5 is a schematic diagram of a storage and computing separation architecture applicable to the data processing system 10;

[0042] FIG6 is a schematic diagram of a storage-computing integrated architecture applicable to the data processing system 10;

[0043] FIG7 is a flow chart of an exemplary data processing method provided by the present application;

[0044] FIG8 is a flowchart of another exemplary data processing method provided in this application. DETAILED DESCRIPTION

[0045] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, various non-limiting embodiments of the embodiments of the present application will be exemplified below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of them. Based on the embodiments in this application, all other embodiments obtained based on the above content are within the scope of protection of this application.

[0046] 1 is a schematic diagram of the structure of an exemplary data processing system 10, which includes a database cluster 100 and a storage device 301. For example, the database cluster 100 may be a database cluster constructed based on a MySQL database, or may be a database cluster constructed based on another type of database.

[0047] As shown in Figure 1, the database cluster 100 includes a service layer and a storage engine layer. The service layer includes multiple service instances, and Figure 1 uses service instance 101 and service instance 102 as an example for illustration. The storage engine layer includes at least one storage engine instance, and Figure 1 uses storage engine instance 201 as an example for illustration. Exemplarily, each service instance in the database cluster 100 can be implemented through one or more processes, and each storage engine instance can be implemented through one or more processes, and the service instance and the storage engine instance are implemented through different processes. In actual application, the service instances and storage engine instances in the database cluster 100 can also be implemented through other methods, which are not limited to this.

[0048] Service instance 101 and service instance 102 are connected to storage engine instance 201 respectively. Taking the connection between service instance 101 and storage engine instance 201 as an example, when service instance 101 and storage engine instance 201 are both implemented through processes, service instance 101 and storage engine instance 201 can establish a communication connection through shared memory, that is, service instance 101 can write data into shared memory, and storage engine instance 201 can obtain data from the shared memory, thereby realizing the transfer of data between the two processes. Alternatively, service instance 101 and storage engine instance 201 can also be connected through a communication network. Similarly, a communication connection can also be established between service instance 102 and storage engine instance 201 through a network or through shared memory. Storage engine instance 201 can provide data reading and writing services for multiple service instances connected to it (such as service instance 101 and service instance 102 in Figure 1).

[0049] The storage engine instance 201 can access the persistent storage layer, such as accessing the persistent storage layer through the network. As shown in Figure 1, the persistent storage layer includes a storage device 301. The storage device 301 not only persistently stores data, but also persistently stores the metadata corresponding to the data. Among them, metadata refers to data used to describe data, such as attribute information such as the storage location, data security permissions, and data type of the data. Exemplarily, the storage device 301 can store metadata in the form of a system table, which can be used to store metadata of various data objects. In the data processing system 10 shown in Figure 1, the data and metadata stored in the storage device 301 can be accessed by the storage engine instance 201.

[0050] It should be noted that in the data processing system 10 shown in Figure 1, only one copy of metadata can be persistently stored in the storage device 301, and there is no need to persistently store a copy of the metadata in the storage device 301 in the local storage area of ​​the service instance 101 and the service instance 102 respectively.

[0051] After the service instance 101 obtains a database statement indicating an update to the data 1 in the storage device 301 , the service instance 101 may parse the database statement, generate at least one operation command for updating the data 1 , and send the at least one operation command to the storage engine instance 201 .

[0052] After receiving the at least one operation command, the storage engine instance 201 can obtain the metadata 1 corresponding to the data 1 from the storage device 301, and update the data 1 according to the metadata 1 and the received operation command, generate data 2 (that is, generate new data), and further update the metadata 1 corresponding to the data 1 to the metadata 2 corresponding to the data 2.

[0053] Since metadata 1 is stored in the storage device 301 and the storage engine instance 201 can access the metadata 1 in the storage device 301, the service instance 101 and the service instance 102 can share the same metadata in the storage device 301 through the storage engine instance 201. This allows the service instance 101 to update the metadata 1 corresponding to the data 1 in the storage device 301 to the metadata 2 corresponding to the data 2 through the storage engine instance 201 without having to perform the metadata synchronization process, that is, it is not necessary to synchronize metadata 2 to the local storage area (for persistent storage) of the service instance 101 and the service instance 102, thereby avoiding the resource consumption caused by synchronizing metadata in the database cluster 100, and the database cluster 100 can use more resources to process business, thereby improving the performance of the database cluster 100.

[0054] Furthermore, service instance 101 and service instance 102 can share the same metadata in storage device 301 through storage engine instance 201. This eliminates the need for service instance 101 and service instance 102 to persistently store a separate copy of metadata in a local storage area, thereby avoiding the storage resource consumption caused by persistently storing multiple copies of metadata in database cluster 100. This allows database cluster 100 to utilize more storage resources to process services, thereby further improving the performance of database cluster 100.

[0055] In addition, during the initialization process of the database cluster 100 and the storage device 301, only one service instance in the database cluster 100 can create a system table for storing metadata in the storage device 301 to complete the initialization of the storage device 301, thereby simplifying the initialization process of the database cluster 100.

[0056] In addition to updating the storage device 301 , the storage engine instance 201 can also read data in the storage device 301 and provide it to the service instance 101 .

[0057] It should be noted that the data processing system 10 shown in FIG1 is merely an example and is not intended to be limiting.

[0058] As a first other implementation example, referring to FIG2 , based on the data processing system 10 shown in FIG1 , the service layer may further include a service instance 103, and the storage engine layer may further include a storage engine instance 202, and the storage engine instance 202 is connected to the service instance 103 and the storage device 301 respectively, and provides data reading and writing services to the service instance 103 by accessing the storage device 301. Among them, the storage engine instance 201 and the storage engine instance 202 can be storage engine instances of the same type, or can be storage engine instances of different types. Moreover, the storage engine instance 201 and the storage engine instance 202 can share access to the data in the storage device 301 in the persistent storage layer. In actual application, the storage device 301 can also be referred to as shared storage. In this way, after updating the metadata in the storage device 301, the network resources required to send the updated metadata from the storage engine instance 201 to the storage engine instance 202 can be avoided.

[0059] As a second other implementation example, referring to FIG3 , based on the data processing system 10 shown in FIG1 , the storage device 301 in the persistent storage layer may include multiple sub-storage devices, and FIG3 is taken as an example to illustrate the sub-storage device 3011 and the sub-storage device 3012. Each sub-storage device is used to store part of the data in the storage device 301 and the metadata corresponding to the part of the data, and different sub-storage devices are used to store different parts of the data and metadata. In actual application scenarios, the data in the storage device 301 can be divided into multiple shards, and each sub-storage device can store one or more shards. For example, based on the primary keyword of the data in the storage device 301, the value range of the primary keyword can be divided into multiple intervals, so that each sub-storage device can store data with the primary keyword located in one or more intervals, and the keywords of the data stored in different sub-storage devices are located in different intervals.

[0060] Furthermore, based on the database cluster 100 shown in FIG1 , the storage engine layer may further include a storage engine instance 202, and the storage engine instance 202 is connected to the service instance 101 and the service instance 102, respectively. In this case, the storage engine instance 201 is specifically connected to the sub-storage device 3011, and the storage engine instance 202 is specifically connected to the sub-storage device 3012. Therefore, the storage engine instance 201 and the storage engine instance 202 are respectively responsible for accessing different sub-storage devices. Specifically, the storage engine instance 201 may be responsible for reading and writing data and metadata in the sub-storage device 3011; the storage engine instance 202 may be responsible for reading and writing data and metadata in the sub-storage device 3012. In this way, the service instance 101 and the service instance 102 can access the shard data A and the metadata corresponding to the shard data A in the sub-storage device 3011 through the storage engine instance 201, and access the shard data B and the metadata corresponding to the shard data B stored in the sub-storage device 3012 through the storage engine instance 202.

[0061] As a third alternative implementation example, see Figure 4. Based on the data processing system 10 described in Figure 1, the storage device 301 in the persistent storage layer may include a metadata storage device and multiple data storage devices. Figure 3 illustrates this by taking metadata storage device 301-1, data storage device 301-2, and data storage device 301-3 as an example. Each data storage device is used to store a portion of the data in device 301, with different data storage devices storing different portions of data. Metadata storage device 301-1 is used to store metadata corresponding to all data. That is, metadata corresponding to the data in each data storage device can be centrally stored in metadata storage device 301-1.

[0062] Furthermore, based on the database cluster 100 shown in FIG1 , the storage engine layer may further include a storage engine instance 202, and the storage engine instance 202 is connected to the service instance 101 and the service instance 102, respectively. In this case, the storage engine instance 201 is connected to the metadata storage device 301-1 and the data storage device 301-2, respectively, and the storage engine instance 202 is connected to the metadata storage device 301-1 and the data storage device 301-3, respectively. The storage engine instance 201 and the storage engine instance 202 are responsible for accessing different data storage devices, and both can access the metadata storage device 301-1.

[0063] This application does not limit the specific architecture of the data processing system 10. For example, in actual application, the database clusters 100 shown in Figures 1 to 4 can be combined to obtain a new data processing system.

[0064] In actual application, in the first implementation method, the data processing system 10 shown in Figures 1 to 4 above can be applied to the device cluster adopting the storage and computing separation architecture shown in Figure 5.

[0065] In the storage and computing separation architecture shown in FIG5 , multiple computing nodes 410 and multiple storage nodes 400 may be included.

[0066] Among them, a service instance and a storage engine instance can be deployed on each computing node 410. For example, the service instance 101 and the storage engine instance 201 in Figure 1 can be deployed on one computing node 410, and the service instance 102 can be deployed on another computing node 410. In addition, different computing nodes 410 can communicate with each other, so that the network communication between the service instance 101 and the service instance 102 can be completed through the network communication between the computing nodes 410 and the computing nodes 410. Alternatively, the service instance 101, the service instance 102 and the storage engine instance 201 in Figure 1 can also be all deployed on the same computing node 410. The storage device 301 in Figure 1 can be implemented by one or more storage nodes 400.

[0067] Each computing node 410 is a computing device including a processor, such as a server, a desktop computer, etc., and the processor can be a central processing unit (CPU). In terms of hardware, as shown in Figure 5, the computing node 410 includes at least a CPU 412, a memory 413, and a network card 414. Among them, the CPU 412 is used to process database statements from outside the computing node 410, or database statements generated inside the computing node 410. The CPU 412 reads data from the memory 413, or when the total amount of data in the memory 413 reaches a certain threshold, the CPU 412 sends the data stored in the memory 413 to the storage node 400 for persistent storage. Figure 5 only shows one CPU 412. In actual applications, there are often multiple CPUs 412, and one CPU 412 has one or more CPU cores. This embodiment does not limit the number of CPUs or the number of CPU cores.

[0068] Memory 413 refers to internal storage that directly exchanges data with the processor. It can read and write data at any time and at high speed, and serves as temporary data storage for the operating system or other running programs. Memory includes at least two types of memory. For example, memory can be either random access memory or read-only memory (ROM). In practical applications, multiple memories 413, as well as different types of memories 413, can be configured in computing node 410. This embodiment does not limit the number and type of memory 413.

[0069] Network card 414 is used to communicate with storage node 400 or other computing nodes 410. For example, when the total amount of data in memory 413 reaches a certain threshold, computing node 410 can send a request to storage node 400 via network card 414 to persistently store the data. Furthermore, computing node 410 may also include a bus for communication between components within computing node 410. In practical applications, computing node 410 may also have a small number of internal hard drives or external hard drives.

[0070] Each computing node 410 can access the storage node 400 through the network. Each storage node 400 may include a controller 401, a network card 404 and a hard disk 405, and the number of controllers 401, network cards 404 and hard disks 405 can be any number. The network card 404 is used to communicate with the computing node 410, or can communicate with other storage nodes 400. The hard disk 405 is used for persistent storage of data, and can be a disk or other types of storage media, such as a solid-state drive or a shingled magnetic recording hard disk. The controller 401 is used to convert the address carried in the read / write data request into an address that the hard disk can recognize according to the read / write data request sent by the computing node 410, and write data to the hard disk 405 or read data from the hard disk 405 according to the address.

[0071] In a second implementation, the data processing system 10 shown in FIG. 1 to FIG. 4 can be applied to a device cluster with a storage-computing integrated architecture as shown in FIG. 6 .

[0072] In the storage and computing integrated architecture shown in FIG6 , multiple servers 410 may be included.

[0073] Among them, a service instance and a storage engine instance can be deployed on each server 410. For example, the service instance 101 and the storage engine instance 201 in Figure 1 can be deployed on one server 410, and the service instance 102 can be deployed on another server 410. In addition, the servers 410 can communicate with each other, so that the network communication between the service instance 101 and the service instance 102 can be completed through the network communication between the servers 410 and the servers 410. Alternatively, the service instance 101, the service instance 102 and the storage engine instance 201 in Figure 1 can also be all deployed on the same server 410. The storage device 301 in Figure 1 can be implemented by the hard disk 405 in one or more servers 410.

[0074] Each server 410 can be a device with computing and storage capabilities. In terms of hardware, as shown in FIG6 , server 410 includes at least a CPU 412 , memory 413 , a network interface card 414 , and a hard disk 405 . The specific implementation of CPU 412 , memory 413 , network interface card 414 , and hard disk 405 can be found in the above description of CPU 412 , memory 413 , network interface card 414 , and hard disk 405 in the architecture shown in FIG5 , and is not further described here.

[0075] To facilitate understanding and explanation, the following describes in detail the process of the database cluster 100 updating the data and metadata in the storage device 301 based on the data processing system 10 shown in FIG. 1 .

[0076] Typically, during operation, service instance 101 and service instance 102 may obtain database statements, such as receiving database statements sent by a user through a client, or automatically generating database statements during the execution of a business by service instance 101 and service instance 102. For example, a database statement may be a structured query language (SQL) statement, specifically a data definition language (DDL) statement within an SQL statement; or, the database statement may be other types of statements, which are not limited in this embodiment. For ease of understanding, this embodiment uses the example of service instance 101 obtaining a database statement as an example, and the database statement is used to instruct an update of data 1 in storage device 301.

[0077] After obtaining the database statement, the service instance 101 can obtain metadata from the storage device 301 through the storage engine instance 201 and, based on the metadata, perform lexical analysis, syntax analysis, and semantic checking on the database statement to determine whether the database statement is legal and generate a corresponding syntax analysis tree. When the database statement is determined to be legal, the service instance 101 can use the optimizer to optimize the syntax analysis tree and generate an execution plan tree corresponding to the database statement. The execution plan tree can define the operations to be performed sequentially during the process of updating the data 1 in the storage device 301. Then, based on the execution plan tree, the service instance 101 can generate one or more operation commands, which are used to instruct the storage engine instance 201 to perform specific operations when updating the data 1 in the storage device 301 and the metadata 1 corresponding to the data 1.

[0078] In actual applications, before obtaining database statements, service instance 101 can pre-load metadata stored in storage device 301. For example, metadata from storage device 301 can be stored in a cache area of ​​service instance 101 via a data dictionary to improve the efficiency of service instance 101 in parsing database statements. The data dictionary is a directory for recording metadata, and the metadata contained in the data dictionary can be loaded from storage device 301.

[0079] In the database cluster 100, since multiple service instances can read any data or modify any data in the storage device 301 through the storage engine instance 201, when a service instance updates the data in the storage device 301 through the storage engine instance, the service instance can set a lock for the metadata corresponding to the data and notify the remaining service instances to also set a lock for the metadata to prohibit the remaining service instances from updating (and reading) the data described by the metadata, thereby ensuring that multiple service instances will not concurrently modify the same data and the same metadata.

[0080] As an implementation example, during the process of generating an operation command (or before generating the operation command), service instance 101 sets a lock for metadata 1 and notifies storage engine instance 201 of the lock. For example, based on the metadata lock (MDL) mechanism, service instance 101 can set the value of a lock variable (e.g., MDL_lock) corresponding to metadata 1 to a write lock state, and write the lock variable to shared memory between service instance 101 and storage engine instance 201. Storage engine instance 201 can then determine that a lock has been set on metadata 1 based on the value of the lock variable in the shared memory. Storage engine instance 201 can then notify service instance 102 via shared memory that a lock has been set on its cached metadata 1. In this way, service instance 102 can set a lock on its cached metadata 1, thereby prohibiting service instance 102 from updating or reading data 1 described by the metadata 1 while the metadata 1 is locked. Storage engine instance 201 can establish different shared memory channels for service instance 101 and service instance 102, allowing communication with different service instances based on different shared memory channels.

[0081] Furthermore, when database cluster 100 further includes storage engine instance 202 and service instance 103 as shown in FIG2 , storage engine instance 201, after determining that a lock has been set on metadata 1, may further generate a lock message and send the lock message to storage engine instance 202 via the communication network. The lock message includes metadata 1. Based on the lock message, storage engine instance 202 may notify service instance 103 to set a lock on metadata 1 via shared memory between the storage engine instance 202 and service instance 103.

[0082] After completing the lock setting, service instance 102 (and other service instances) can notify service instance 101 through storage engine instance 201 (and other storage engine instances) that the lock setting for metadata 1 has been completed. Specifically, it can be to feedback a response of the setting completion to service instance 101. In this way, service instance 101 can send the at least one generated operation command to storage engine instance 201 in sequence. Exemplarily, the operation command may include keywords for indicating the operation, such as "CREATE", "ALTER", or "DROP", etc., which respectively represent creating, modifying, deleting data in storage device 301; the operation command may also include relevant information of the object being operated (i.e., data 1).

[0083] In this embodiment, the storage engine instance 201 can sequentially execute the received operation commands and perform corresponding update operations on the data 1 in the storage device 301, so as to update the data 1 stored in the storage device 301 to data 2. For example, the storage engine instance 201 can create a new table in the storage device 301 and use the new table to store data; or the storage engine instance 201 can modify the data in some tables in the storage device 301; or the storage engine instance 201 can delete some tables in the storage device 301, or delete some data in some tables, etc. Accordingly, in the process of executing the operation commands, the storage engine instance 201 will also generate metadata 2 for the data 2 and update the metadata 1 in the storage device 301 to metadata 2, so as to achieve synchronous update of the data and metadata in the storage device 301.

[0084] After executing the operation command, the storage engine instance 201 may feed back a response to the service instance 101 to notify the service instance 101 that the operation command is completed.

[0085] Since the metadata 1 in the storage device 301 has been updated to metadata 2, that is, the metadata 1 cached in each service instance has expired, therefore, after receiving the execution completion response feedback from the storage engine 201, the service instance 101 can invalidate the currently cached metadata 1, such as the service instance 101 can invalidate the cached data dictionary (including the metadata 1) to avoid the service instance 101 subsequently executing an erroneous data processing process based on the expired metadata 1.

[0086] Furthermore, the service instance 101 may also notify other service instances to invalidate the metadata 1 cached therein.

[0087] In a specific implementation, service instance 101 may generate an invalidation message that may include metadata 1, indicating that metadata 1 should be invalidated. Service instance 101 may then provide the invalidation message to storage engine instance 201 via shared memory. Storage engine instance 201 may then send the invalidation message to service instance 102 via shared memory. In this way, service instance 102 may invalidate its cached metadata 1 based on the invalidation message in shared memory.

[0088] Furthermore, when the database cluster 100 further includes the storage engine instance 202 and the service instance 103 as shown in FIG2 , the storage engine instance 201 can send the invalidation message to the storage engine instance 202 via the network, so that the storage engine instance 202 can provide the invalidation message to the service instance 103 via shared memory. In this way, the service instance 103 can invalidate its cached metadata 1 based on the invalidation message. In actual application, when the storage engine instance 201 or the storage engine instance 202 is connected to multiple service instances, the storage engine instance 201 or the storage engine instance 202 can notify all service instances connected to it (except the service instance 101) to invalidate their cached metadata 1.

[0089] The above example uses the service instance 101 notifying other service instances to invalidate the cached metadata 1. In other embodiments, the storage engine instance 201 may also proactively notify other service instances to invalidate their cached metadata 1 after executing the completed operation command. This is not limited to this.

[0090] After completing the invalidation of metadata 1, service instance 102 (and other service instances) can feedback an invalidation completion response to service instance 101 through storage engine instance 202 and storage engine instance 201 to notify service instance 101 that the invalidation of metadata 1 is completed.

[0091] In actual application, after invalidating the data dictionary (including metadata 1), service instances 101 and 102 in database cluster 100 can access the latest version of metadata stored in storage device 301, namely metadata 2, through storage engine instance 201. Service instances 101 and 102 can then load metadata 2 to generate a new data dictionary and cache it. In this way, different service instances can load the same metadata to generate a data dictionary, thereby ensuring the consistency of the metadata cached by different service instances.

[0092] Because each service instance has a lock corresponding to metadata 1, service instance 101 can release the lock set for metadata 1 after receiving the invalidation completion response, such as by setting the value of the lock variable corresponding to metadata 1 to unlocked. Simultaneously, service instance 101 can notify storage engine instance 201 via shared memory that metadata 1 has been unlocked. In this way, storage engine instance 201 can notify service instance 102 via shared memory with service instance 102 to release the lock set for metadata 1.

[0093] Furthermore, when the database cluster 100 may further include a storage engine instance 202 and a service instance 103 as shown in FIG2 , the storage engine instance 201 may generate an unlock message and send the unlock message to the storage engine instance 202 via the communication network. The unlock message includes metadata 1. After receiving the unlock message, the storage engine instance 202 may provide the unlock message to the service instance 103 via the shared memory between the storage engine instance 202 and the service instance 103, so that the service instance 103 may release the lock set for metadata 1 according to the unlock message. In actual application, when the storage engine instance 201 and the storage engine instance 202 are connected to multiple service instances, the storage engine instance 201 and the storage engine instance 202 may also notify the multiple service instances (except the service instance 101) via the shared memory to release the lock set for metadata 1.

[0094] In a further possible implementation, since the storage engine instance 201 may fail during the process of updating the data and metadata in the storage device 301, the service instance 101 can ensure the atomicity of updating the storage device 301 by committing a transaction before invalidating the metadata 1 and releasing the lock set for the metadata 1.

[0095] In the data processing system 10 shown in FIG2 , after determining that the storage engine instance 201 has completed the execution of the operation command (before invalidating metadata 1), the service instance 101 can send a transaction to the storage engine instance 201. The transaction records the update operations for data 1 and metadata 1. For example, the transaction includes the database statement obtained by the service instance 101, so that the storage engine instance 201 performs the operation of committing the transaction, such as marking the database statement as "commit" in the log to indicate that the transaction has been completed. Thus, when the storage engine instance 201 feedbacks that the transaction is successfully committed, the service instance 101 executes the corresponding process of invalidating metadata 1 and unlocking metadata 1. If the storage engine instance 201 feedbacks that the transaction fails to commit to the service instance 101, the service instance 101 can instruct the storage engine instance 201 to roll back the update operations for the data 1 and metadata 1 to restore the storage device 301 to the state before the data update.

[0096] Among them, if the storage engine instance 201 fails before successfully submitting the transaction, the storage engine instance 202 can take over the subsequent processing flow for the database statement. Specifically, since the transaction was not submitted successfully, the storage engine instance 202 will roll back the update operation of the storage engine instance 201 on the data 1 and metadata 1 in the storage device 301, that is, restore the data 2 and metadata 2 in the storage device 301 to data 1 and metadata 1. In addition, since the lock corresponding to metadata 1 still exists in the service instance 103, the storage engine instance 202 can notify the service instance 103 through shared memory to release the lock set for metadata 1, so that the metadata 1 cached in the service instance 103 remains valid, that is, the subsequent service instance 103 can still access the data 1 in the storage device 301 through the metadata 1. In this way, through the takeover of the storage engine instance 202, fault recovery for the storage engine instance 201 can be achieved.

[0097] If storage engine instance 201 fails after successfully committing a transaction, the update operation on storage device 301 takes effect, that is, the data and metadata stored in storage device 301 are updated data 2 and metadata 2. Then, when storage engine instance 202 takes over the subsequent processing flow of the database statement, it can specifically notify service instance 103 through shared memory to invalidate the cached metadata 1 and notify service instance 103 to release the lock set for metadata 1, thereby achieving failure recovery for storage engine instance 201.

[0098] It should be noted that this embodiment is illustrated by taking the service instance 101 updating the data 1 and metadata 1 in the storage device 301 as an example. The implementation process of the service instance 101 updating other data and other metadata in the storage device 301, as well as the implementation process of the service instance 102 updating other data and other metadata in the storage device 301 (and other databases), can refer to the above-mentioned relevant descriptions and will not be repeated here.

[0099] It is worth noting that the above description uses the data processing system 10 shown in FIG1 as an example to describe the process by which database cluster 100 updates data and metadata in storage device 301. The specific implementation process of service instance 101 updating data and metadata in storage device 301 in data processing system 10 shown in FIG2 through FIG4 is similar to the implementation process of updating data and metadata in storage device 301 in data processing system 10 shown in FIG1 . For further understanding, please refer to the relevant descriptions of the above embodiments and will not be repeated here.

[0100] For ease of understanding, an embodiment of the data processing method provided in this application is described below in conjunction with the accompanying drawings.

[0101] 7 is a flow chart illustrating a data processing method according to an embodiment of the present application. This method can be applied to the data processing system 10 shown in FIG. 1 or FIG. 2 , or to other applicable data processing systems. For ease of explanation, this embodiment uses the data processing system 10 shown in FIG. 2 as an example to describe the process of updating data and metadata in the storage device 301 by the service instance 101 through the storage engine instance 101.

[0102] The data processing method shown in FIG7 may specifically include:

[0103] S701 : The service instance 101 obtains a database statement, which is used to instruct to update the data 1 in the storage device 301 .

[0104] Exemplarily, the database statement may be, for example, a DDL statement, etc. The database statement may be sent by a user to the service instance 101 through a client or other device, or the service instance 101 may generate the corresponding database statement during operation.

[0105] S702: The service instance 101 parses the database statement according to the cached data dictionary and generates an operation command.

[0106] In this embodiment, service instance 101 (as well as service instance 103 and service instance 102) can pre-load metadata from storage device 301 and generate a data dictionary in a local cache area based on the loaded metadata, so as to improve the efficiency of parsing database statements by using the cached data dictionary.

[0107] S703: The service instance 101 sets a lock for the metadata 1 corresponding to the data 1, and notifies the storage engine instance 201 of the lock through the shared memory 1.

[0108] S704 : The storage engine instance 201 notifies the service instance 102 through the shared memory 2 to set a lock for the cached metadata 1 , and sends a lock message to the storage engine instance 202 .

[0109] The locking message may include metadata 1.

[0110] S705: The storage engine instance 202 notifies the service instance 103 through the shared memory 3 to set a lock for the metadata 1.

[0111] S706: Service instance 103 and service instance 102 set a lock for metadata 1 in the cached data dictionary.

[0112] In actual application, after completing the lock setting for metadata 1, service instance 103 and service instance 102 can feedback a response to service instance 101 through storage engine instance 201 and storage engine instance 202 to notify service instance 101 that they have completed the lock setting for metadata 1.

[0113] S707: The service instance 101 sends an operation command to the storage engine instance 201 through the shared memory 1.

[0114] S708: The storage engine instance 201 executes the operation command to update the data 1 and metadata 1 in the storage device 301 to data 2 and metadata 2 respectively.

[0115] After completing the update of the data and metadata, the storage engine instance 201 can feed back a response indicating that the operation command has been executed to the service instance 101 through the shared memory 1.

[0116] S709 : The service instance 101 sends a transaction to the storage engine instance 201 through the shared memory 1 . The transaction includes the database statement obtained by the service instance 101 .

[0117] S710: The storage engine instance 201 commits the transaction.

[0118] In this embodiment, the database cluster 100 supports providing a transaction mechanism to ensure the atomicity of the update operation of each service instance on the storage device 301.

[0119] After successfully submitting the transaction, the storage engine instance 201 may feed back a response indicating successful transaction submission to the service instance 101 through the shared memory 1 .

[0120] S711 : The service instance 101 invalidates the cached data dictionary, which includes metadata 1 , and notifies the storage engine instance 201 via the shared memory 1 .

[0121] S712 : The storage engine instance 201 notifies the service instance 102 through the shared memory 2 to invalidate the cached data dictionary, and sends an invalidation message to the storage engine instance 202 .

[0122] S713: The storage engine instance 202 notifies the service instance 103 through the shared memory 3 to invalidate the cached data dictionary.

[0123] S714: Service instance 103 and service instance 102 execute the operation of invalidating the data dictionary.

[0124] After completing the invalidation of the data dictionary, service instance 103 and service instance 102 may feedback a response to service instance 101 through storage engine instance 201 and storage engine instance 202 to notify service instance 101 that they have completed the invalidation of the data dictionary.

[0125] S715: The service instance 101 releases the lock set for the metadata 1 corresponding to the data 1, and notifies the storage engine instance 201 of the lock through the shared memory 1.

[0126] S716: The storage engine instance 201 notifies the service instance 102 through the shared memory 2 to release the lock set for the cached metadata 1, and sends an unlock message to the storage engine instance 202.

[0127] The unlocking message may include metadata 1.

[0128] S717: The storage engine instance 202 notifies the service instance 103 through the shared memory 3 to release the lock set for the metadata 1.

[0129] S718: Service instance 103 and service instance 102 release the lock set for metadata 1.

[0130] Since metadata 1 is stored in the storage device 301, and the storage engine instance 201 and the storage engine instance 202 can both access the metadata 1 in the storage device 301, the service instance 101, the service instance 103 and the service instance 102 can share the same metadata in the storage device 301 through the storage engine instance 201 and the storage engine instance 202. This allows the service instance 101 to update the metadata 1 in the storage device 301 to metadata 2 through the storage engine instance 201 without having to synchronize the updated metadata 2 to the local storage areas of the service instance 103 and the service instance 102, thereby avoiding the resource consumption caused by synchronizing metadata in the database cluster 100, and the database cluster 100 can process business based on more resources, thereby improving the performance of the database cluster 100.

[0131] Furthermore, service instances 101, 103, and 102 can share the same metadata in storage device 301 through storage engine instance 201. This eliminates the need for database cluster 100 to persistently store a separate copy of metadata locally on service instances 101, 103, and 102, as is required for a standalone database. This avoids the storage resource consumption associated with persistently storing multiple copies of metadata within database cluster 100. This allows database cluster 100 to process services using more storage resources, further improving its performance.

[0132] In addition, the database cluster 100 supports a transaction mechanism, so that the database cluster 100 can complete the modification of data and metadata in the storage device 301 within one transaction, which can ensure the atomicity of the modification of data and metadata.

[0133] Furthermore, after completing the update of metadata 1 in storage device 301, each service instance will invalidate the data dictionary (including metadata 1) and can generate a new data dictionary (including metadata 2) by loading the same metadata from storage device 301, thereby ensuring that the metadata cached by different service instances remain consistent.

[0134] Similarly, the service instance 103 may also update the data and metadata in the storage device 301 through the storage engine instance 202 .

[0135] Specifically, taking the example of service instance 103 continuing to update data 2 and metadata 2 in storage device 301, service instance 103 may generate a new database statement during operation, which is used to update data 2 and metadata 2 in storage device 301. Then, service instance 101 may parse the new database statement based on the cached data dictionary (e.g., the data dictionary may be generated by reloading metadata from the persistent storage layer), generate a new operation command, and send the new operation command to storage engine instance 202. Based on the received operation command, storage engine instance 202 updates data 2 in storage device 301 to data 3, and updates metadata 2 in storage device 301 to metadata 3, thereby updating the data and metadata in storage device 301.

[0136] In actual application, in the process of updating data 2 and metadata 2, you can also refer to the method flow shown in Figure 7 above to execute processes such as setting a lock for metadata 2, invalidating the data dictionary, committing the transaction, and releasing the lock set for metadata 2. For details, please refer to the relevant description of the embodiment shown in Figure 7 above, which will not be repeated here.

[0137] It should be noted that the method steps in the embodiment shown in FIG7 are merely exemplary and not limiting. In other embodiments, the execution order of the steps may be other, and this is not a limitation. Furthermore, some of the method steps in the embodiment shown in FIG7 may not be executed, such as the steps related to caching the data dictionary and invalidating the data dictionary.

[0138] 8 is a flow chart illustrating a data processing method according to an embodiment of the present application. This method can be applied to the data processing system 10 shown in FIG. 3 or FIG. 4 , or to other applicable data processing systems. For ease of explanation, this embodiment uses the data processing system 10 shown in FIG. 3 as an example to describe the process of updating data and metadata in the sub-storage device 3011 by the service instance 101 through the storage engine instance 101.

[0139] The data processing method shown in FIG8 may specifically include:

[0140] S801 : The service instance 101 obtains a database statement, which is used to instruct to update the data 1 in the sub-storage device 3011 .

[0141] S802: The service instance 101 parses the database statement according to the cached data dictionary and generates an operation command.

[0142] S803: The service instance 101 sets a lock for the metadata 1 corresponding to the data 1, and sends a notification message 1 to the storage engine instance 201 through the network to notify that the metadata 1 has been locked.

[0143] S804: The storage engine instance 201 sends a lock message to the service instance 103 and the service instance 102 respectively. The lock message includes metadata 1.

[0144] S805: Service instance 103 and service instance 102 respectively set a lock for metadata 1 in the cached data dictionary.

[0145] After completing setting the lock for metadata 1 , service instance 103 and service instance 102 may respectively feedback a response to service instance 101 through storage engine instance 201 to notify service instance 101 that they have completed setting the lock for metadata 1 .

[0146] S806: The service instance 101 sends the operation command generated based on the database statement to the storage engine instance 201 through the network.

[0147] S807: The storage engine instance 201 executes the operation command to update the data 1 and metadata 1 in the sub-storage device 3011 to data 2 and metadata 2 respectively.

[0148] After completing the update of data and metadata, the storage engine instance 201 can feed back a response indicating that the operation command has been executed to the service instance 101 through the network.

[0149] S808: The service instance 101 sends a transaction to the storage engine instance 201 through the network. The transaction includes the database statement obtained by the service instance 101.

[0150] S809: The storage engine instance 201 commits the transaction.

[0151] After successfully submitting the transaction, the storage engine instance 201 may feed back a response indicating successful transaction submission to the service instance 101 through the shared memory 1 .

[0152] S810: The service instance 101 invalidates the cached data dictionary, which includes metadata 1, and sends a notification message 2 to the storage engine instance 201 via the network to notify that the metadata 1 has been invalidated.

[0153] S811: The storage engine instance 201 sends invalidation messages to the service instance 103 and the service instance 102 respectively.

[0154] S812: Service instance 103 and service instance 102 invalidate their cached data dictionaries respectively.

[0155] After completing the invalidation of the data dictionary, the service instance 103 and the service instance 102 may respectively feedback a response to the service instance 101 through the storage engine instance 201 to notify the service instance 101 that they have completed the invalidation of the data dictionary.

[0156] S813: The service instance 101 releases the lock set for the metadata 1 and sends a notification message 3 to the storage engine instance 201 via the network to notify that the metadata 1 has been unlocked.

[0157] S814: The storage engine instance 201 sends an unlock message to the service instance 103 and the service instance 102 through the network respectively.

[0158] The unlocking message may include metadata 1.

[0159] S815: The service instance 103 and the service instance 102 release the locks set for the metadata 1 respectively.

[0160] Thus, when updating data and metadata in sub-storage device 3011, there's no need to synchronize metadata within database cluster 100. This avoids the resource consumption associated with metadata synchronization and improves the performance of database cluster 100. Furthermore, in data processing system 10, only one copy of the updated metadata can be stored in sub-storage device 3011, eliminating the need to consume storage resources in database cluster 100 to persistently store metadata, thus reducing the amount of persistent storage resources required for metadata storage. Furthermore, database cluster 100 supports a transaction mechanism, allowing modifications to data and metadata in sub-storage device 3011 to be completed within a single transaction, ensuring the atomicity of data and metadata modifications.

[0161] In actual application scenarios, the service instance 101 can not only update the data and metadata in the sub-storage device 3011 through the storage engine instance 201, but also update the data and metadata in the sub-storage device 3012 through the storage engine instance 202.

[0162] Specifically, during operation, service instance 101 may generate a new database statement used to update data 3 and metadata 3 corresponding to data 3 in sub-storage device 3012. Service instance 101 may then parse the new database statement based on a cached data dictionary (e.g., the data dictionary may be generated by reloading metadata from the persistent storage layer), generate a new operation command, and send the new operation command to storage engine instance 202. Based on the received operation command, storage engine instance 202 updates data 3 in sub-storage device 3012 to data 4 and metadata 3 in sub-storage device 3012 to metadata 4, thereby updating the data and metadata in sub-storage device 3012.

[0163] In actual application, in the process of updating data 3 and metadata 3, you can also refer to the method flow shown in Figure 8 above to execute processes such as setting a lock for metadata 3, invalidating the data dictionary, committing the transaction, and releasing the lock set for metadata 3. For details, please refer to the relevant description of the embodiment shown in Figure 8 above, which will not be repeated here.

[0164] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned data processing method.

[0165] The present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the computer program product fully or partially generates the process or function described in the present application.

[0166] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0167] The computer program product may be a software installation package. When any of the aforementioned data processing methods is required, the computer program product may be downloaded and executed on a computing device.

[0168] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0169] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects associated with each other are in an "or" relationship. In the embodiments of the present application. "Simultaneously" means within the same time period, including situations at the same time. The terms "first", "second", etc. in the specification, claims and drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, and this is merely a way of distinguishing objects with the same properties when describing them in the embodiments of the present application.

[0170] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0171] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A database cluster, characterized in that: The system comprises a plurality of service instances and at least one storage engine instance, wherein the at least one storage engine instance is connected to the plurality of service instances and a storage device, the storage device is used to persistently store first data and first metadata corresponding to the first data, a first storage engine instance in the at least one storage engine instance is connected to a first service instance and a second service instance in the plurality of service instances, and the first metadata is shared by the first service instance and the second service instance; The first service instance is used to obtain a first database statement, where the first database statement is used to instruct to update the first data, and send a first operation command to the first storage engine instance according to the first database statement; The first storage engine instance is used to update the first data according to the first operation command to obtain second data, and update the first metadata to second metadata corresponding to the second data.

2. The database cluster according to claim 1, characterized in that: The at least one storage engine instance further includes a second storage engine instance, the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; The third service instance is used to obtain a second database statement, where the second database statement is used to instruct to update the second data, and send a second operation command to the second storage engine instance according to the second database statement; The second storage engine instance is used to update the second data according to the second operation command to obtain third data, and update the second metadata to third metadata corresponding to the third data.

3. The database cluster according to claim 1, characterized in that: The storage device comprises a first sub-storage device and a second sub-storage device, the first sub-storage device is used to store the first data and the first metadata, and the first sub-storage device and the second sub-storage device are used to store different data and different metadata; The at least one storage engine instance further includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second sub-storage device, and the first storage engine instance is connected to the first sub-storage device; The first storage engine instance is used to access the first sub-storage device, and the second storage engine instance is used to access the second sub-storage device.

4. The database cluster according to claim 1, characterized in that: The storage device comprises a first data storage device, a second data storage device and a metadata storage device, the first data storage device is used to store the first data, the first data storage device and the second data storage device are used to store different data, and the metadata storage device is used to store metadata corresponding to the data in the first data storage device and metadata corresponding to the data in the second data storage device; The at least one storage engine instance further includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second data storage device, the first storage engine instance is connected to the first data storage device, and the first storage engine instance and the second storage engine instance are both connected to the metadata storage device; The first storage engine instance is used to access the first data storage device and the metadata storage device, and the second storage engine instance is used to access the second data storage device and the metadata storage device.

5. The database cluster according to any one of claims 1 to 4, characterized in that: The first service instance is further configured to set a lock for the first metadata before sending the first operation command to the first storage engine instance, and release the lock set for the first metadata after the first metadata is updated to the second metadata; The first storage engine instance is further used to instruct the second service instance to set a lock for the first metadata, and after updating the first metadata to the second metadata, instruct the second service instance to release the lock set for the first metadata.

6. The database cluster according to claim 5, characterized in that: The at least one storage engine instance further includes a second storage engine instance, the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; The first storage engine instance is further used to instruct the third service instance to set a lock for the first metadata through the second storage engine instance before updating the first metadata to the second metadata; Furthermore, after the first metadata is updated to the second metadata, the third service instance is instructed, through the second storage engine instance, to release the lock set for the first metadata.

7. The database cluster according to claim 6, characterized in that: Before the first metadata is updated to the second metadata, the first metadata is cached in the first service instance, the second service instance, and the third service instance; The first service instance is further used to invalidate the first metadata cached by the first service instance after the first metadata is updated to the second metadata; The first storage engine instance is further used to instruct the second service instance to invalidate the first metadata cached by the second service instance, and to instruct the third service instance to invalidate the first metadata cached by the third service instance through the second storage engine instance.

8. The database cluster according to claim 7, characterized in that: The second service instance is used to cache the second metadata in the storage device after the cached first metadata is invalidated.

9. The database cluster according to claim 7 or 8, characterized in that: The first service instance is further configured to send a transaction to the first storage engine instance after the first metadata is updated to the second metadata, the transaction including the database statement; The first storage engine instance is also used to commit the transaction.

10. The database cluster according to claim 9, characterized in that: The first storage engine instance fails before committing the transaction; The second storage engine instance is used to roll back the update operation on the first data and the first metadata, and instruct the third service instance to release the lock set for the first metadata.

11. The database cluster according to claim 9, characterized in that: The first storage engine instance fails after the transaction is committed; The second storage engine instance is used to instruct the third service instance to release the lock set for the first metadata and invalidate the first metadata cached by the third service instance.

12. A data processing method, characterized in that: Applied to a database cluster, the database cluster includes multiple service instances and at least one storage engine instance, the at least one storage engine instance is connected to the multiple service instances and a storage device, the storage device is used to persistently store first data and first metadata corresponding to the first data, a first storage engine instance in the at least one storage engine instance is connected to a first service instance and a second service instance in the multiple service instances, and the first metadata is shared by the first service instance and the second service instance; The method comprises: The first service instance obtains a first database statement, where the first database statement is used to instruct to update the first data; The first service instance sends a first operation command to the first storage engine instance according to the first database statement; The first storage engine instance updates the first data according to the first operation command to obtain second data; The first storage engine instance updates the first metadata to second metadata corresponding to the second data.

13. The method according to claim 12, characterized in that The at least one storage engine instance further includes a second storage engine instance, the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; The method further comprises: The third service instance obtains a second database statement, where the second database statement is used to instruct to update the second data; The third service instance sends a second operation command to the second storage engine instance according to the second database statement; The second storage engine instance updates the second data according to the second operation command to obtain third data; The second storage engine instance updates the second metadata to third metadata corresponding to the third data.

14. The method according to claim 12, characterized in that The storage device comprises a first sub-storage device and a second sub-storage device, the first sub-storage device is used to store the first data and the first metadata, and the first sub-storage device and the second sub-storage device are used to store different data and different metadata; The at least one storage engine instance further includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second sub-storage device, and the first storage engine instance is connected to the first sub-storage device; The first storage engine instance is used to access the first sub-storage device, and the second storage engine instance is used to access the second sub-storage device.

15. The method according to claim 12, characterized in that The storage device comprises a first data storage device, a second data storage device and a metadata storage device, the first data storage device is used to store the first data, the first data storage device and the second data storage device are used to store different data, and the metadata storage device is used to store metadata corresponding to the data in the first data storage device and metadata corresponding to the data in the second data storage device; The at least one storage engine instance further includes a second storage engine instance, the second storage engine instance is connected to the first service instance, the second service instance, and the second data storage device, the first storage engine instance is connected to the first data storage device, and the first storage engine instance and the second storage engine instance are both connected to the metadata storage device; The first storage engine instance is used to access the first data storage device and the metadata storage device, and the second storage engine instance is used to access the second data storage device and the metadata storage device.

16. The method according to any one of claims 12 to 15, characterized in that The method further comprises: Before sending the first operation command to the first storage engine instance, the first service instance sets a lock for the first metadata; The first storage engine instance instructs the second service instance to set a lock for the first metadata; After the first metadata is updated to the second metadata, the first service instance releases the lock set for the first metadata; After updating the first metadata to the second metadata, the first storage engine instance instructs the second service instance to release the lock set for the first metadata.

17. The method according to claim 16, characterized in that The at least one storage engine instance further includes a second storage engine instance, the multiple service instances further include a third service instance, and the second storage engine instance is connected to the third service instance and the storage device; The method further comprises: Before updating the first metadata to the second metadata, the first storage engine instance instructs the third service instance to set a lock for the first metadata through the second storage engine instance; After updating the first metadata to the second metadata, the first storage engine instance instructs the third service instance, through the second storage engine instance, to release the lock set for the first metadata.

18. The method according to claim 17, characterized in that Before the first metadata is updated to the second metadata, the first metadata is cached in the first service instance, the second service instance, and the third service instance; The method further comprises: After the first metadata is updated to the second metadata, the first service instance invalidates the first metadata cached by the first service instance; The first storage engine instance instructs the second service instance to invalidate the first metadata cached by the second service instance, and instructs the third service instance to invalidate the first metadata cached by the third service instance through the second storage engine instance.

19. The method according to claim 18, characterized in that The method further comprises: The second service instance caches the second metadata in the storage device after invalidating the cached first metadata.

20. The method according to claim 18 or 19, characterized in that The method further comprises: After the first metadata is updated to the second metadata, the first service instance sends a transaction to the first storage engine instance, where the transaction includes the database statement; The first storage engine instance commits the transaction.

21. The method according to claim 20, characterized in that The first storage engine instance fails before committing the transaction; The method further comprises: The second storage engine instance rolls back the update operation on the first data and the first metadata; The second storage engine instance instructs the third service instance to release the lock set for the first metadata.

22. The method according to claim 20, characterized in that The first storage engine instance fails after the transaction is committed; The method further comprises: The second storage engine instance instructs the third service instance to release the lock set for the first metadata; The second storage engine instance invalidates the first metadata cached by the third service instance.

23. A data processing system, characterized in that: The data processing system comprises a database cluster as claimed in any one of claims 1 to 11 and a storage device as claimed in any one of claims 1 to 11.

24. A device cluster, characterized in that: The device cluster includes at least one computing node and at least one storage node. The at least one computing node is used to implement the database cluster according to any one of claims 1 to 11, and the at least one storage node is used to implement the storage device according to any one of claims 1 to 11.

25. A computer-readable storage medium, characterized in that: The method comprises instructions which, when executed on a computing device, cause the computing device to perform the steps of the method as claimed in any one of claims 12 to 22.

26. A computer program product comprising instructions, characterized in that When it is executed on at least one computing device, the at least one computing device is caused to perform the steps of the method according to any one of claims 12 to 22.

Citation Information

Patent Citations

  • Database cluster, data processing method, data processing system and related equipment

    CN120123421A

  • Metadata synchronization method and device

    CN111737353A

  • Data management method and device, storage medium and electronic equipment

    CN113761294A

  • Database transaction processing method, electronic equipment and device

    CN115878638A

  • System and method for transaction continuity across failures in a scale-out database

    US20220114058A1