Cloud storage system, data storage method, data reading method, device, storage medium and program product
By persisting data objects to cloud storage nodes with matching read and write performance in the cloud storage system, the problem of high storage costs of cloud native databases is solved and the overall access performance of the cloud storage system is improved.
Patent Information
- Application Number
- PCT/IB2025/052317
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-25
AI Technical Summary
Existing cloud-native databases have high storage costs while ensuring read and write performance, which is difficult to further reduce.
Compute nodes are used to persist data objects to cloud storage nodes whose read and write performance matches their hot and cold levels. By flexibly using different types of cloud storage nodes, the utilization rate of each node is improved and storage costs are reduced.
It reduces cloud storage costs and improves the overall access performance of the cloud storage system while ensuring read and write performance.
Smart Images

Figure IB2025052317_25092025_PF_FP_ABST
Abstract
Description
[0001] Technical Field of the Invention: Cloud Storage System, Data Storage and Reading Method, Device, Storage Medium, and Program Product
[0002]
[0001] The present disclosure relates to the field of cloud computing technology, and more particularly to a cloud storage system, data storage and reading method, device, storage medium, and program product.
[0003]
[0002] With the explosive growth of data migration to the cloud and user data, cloud-native databases are becoming increasingly popular. Cloud-native databases typically use a storage-computing separation architecture to achieve independent elasticity of computing resources and storage resources, high service availability, and pay-as-you-go features. Users' persistent data is typically stored in the storage resources of the cloud-native database (such as a distributed file system) to ensure high data availability and scalability. However, cloud-native databases currently face a contradiction between storage costs and read-write performance. If, while ensuring read-write performance, further reducing cloud storage costs is a technical problem that urgently needs to be solved. SUMMARY
[0004]
[0003] Various aspects of the present disclosure provide a cloud storage system, a data storage and reading method, a device, a storage medium, and a program product to further reduce cloud storage costs while ensuring read and write performance.
[0005]
[0004] An embodiment of the present disclosure provides a cloud storage system, comprising: a computing layer and a storage layer, the computing layer comprising at least one computing node, the storage layer comprising at least two types of cloud storage nodes, and the read and write performance of different types of cloud storage nodes being different; the computing node being used to determine a first data object requiring persistent storage and its attribute information, the attribute information including the hotness or coldness of the first data object; selecting a first cloud storage node from the at least two types of cloud storage nodes, the read and write performance of which is compatible with the hotness or coldness of the first data object, and writing the first data object into the first cloud storage node.
[0006]
[0005] An embodiment of the present disclosure also provides a data reading method, which is applied to a cloud storage system. The cloud storage system includes at least two types of cloud storage nodes, and the read and write performance of different types of cloud storage nodes are different. The method includes: determining a first data object and its attribute information that need to be persistently stored in the cloud storage system, the attribute information including the hotness or coldness of the first data object; selecting a first cloud storage node whose read and write performance is compatible with the hotness or coldness of the first data object from the at least two types of cloud storage nodes, and writing the first data object to the first cloud storage node.
[0007]
[0006] An embodiment of the present disclosure also provides a data reading method, which is applied to a cloud storage system, wherein the cloud storage system includes at least two types of cloud storage nodes, and different types of cloud storage nodes have different read and write performances. The method includes: receiving a read request, the read request being used to request reading a data object; determining a data page corresponding to the data object based on an index tree of the cloud storage system; obtaining metadata information corresponding to the data page, the metadata information including a storage identifier pointing to a storage node where the data page is located and a storage path of the data page on the storage node where the data page is located; reading the data object from the storage node pointed to by the storage identifier in the metadata information based on the storage path in the metadata information, and providing the data object to the initiator of the read request; wherein the storage node where the data page is located is any type of cloud storage node.
[0008]
[0007] An embodiment of the present disclosure further provides a computer device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to perform the steps in the data access method.
[0009]
[0008] The embodiment of the present disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the data access method.
[0010]
[0009] The embodiment of the present disclosure also provides a computer program product, including a computer program / instruction, which, when executed by a processor, enables the processor to implement the steps in the data access method.
[0011] In an embodiment of the present disclosure, a cloud storage system includes at least one computing node and multiple cloud storage nodes with different read / write performance. The computing node can persist data objects that require persistent storage to a cloud storage node whose read / write performance matches the hotness or coldness of the data object. In this way, each cloud storage node can be flexibly used according to the hotness or coldness of the data, thereby increasing the utilization rate of each cloud storage node. This not only meets user application needs, but also reduces storage costs and effectively improves the overall access performance of the cloud storage system.
[0012]
[0011] Furthermore, in the embodiments of the present disclosure, data objects between cloud storage nodes are supported for migration. By transferring data objects from cloud storage nodes with high read / write performance to cloud storage nodes with low read / write performance, it is beneficial to improve the utilization rate of each cloud storage node and reduce storage costs.
[0013]
[0012] The drawings described herein are intended to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are intended to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0014]
[0013] FIG1 is a schematic diagram of the structure of a cloud storage system provided in an embodiment of the present disclosure;
[0015] FIG. 2 is a schematic diagram showing the relationship between the read / write performance and storage cost of different storage media;
[0016]
[0015] FIG3 is a schematic diagram showing the relationship between read / write performance, storage cost, and storage capacity of different storage media;
[0017] FIG4 is a schematic diagram of the structure of a cloud native database provided in an embodiment of the present disclosure;
[0018]
[0017] FIG5 is a flow chart of a data storage method provided in an embodiment of the present disclosure;
[0019] FIG6 is a flow chart of another data reading method provided in an embodiment of the present disclosure;
[0020] FIG. 7 is a schematic diagram of the structure of a data access device provided in an embodiment of the present disclosure;
[0021]
[0020] FIG8 is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure.
[0022] To make the objectives, technical solutions, and advantages of the present disclosure more clearly apparent, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, and are not intended to be exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are intended to fall within the scope of protection of the present disclosure.
[0023]
[0022] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. In addition, the various models involved in this disclosure (including but not limited to language models or large models) comply with relevant laws and standards.
[0024] With the rapid migration of data to the cloud and the explosive growth of user data, cloud-native databases are becoming increasingly popular. Cloud-native databases typically employ a storage-computing separation architecture to achieve independent elasticity of computing and storage resources, high service availability, and pay-as-you-go features. Users' persistent data is typically stored in the cloud-native database's storage resources (such as a distributed file system) to ensure high data availability and scalability. However, cloud-native databases currently face a conflict between storage costs and read / write performance. Further reducing cloud storage costs while ensuring read / write performance is a technical issue that urgently needs to be addressed.
[0025] To this end, embodiments of the present disclosure provide a cloud storage system, data storage and reading method, device, storage medium, and program product. The cloud storage system includes at least one computing node and multiple cloud storage nodes with different read / write performance. The computing node can persist data objects requiring persistent storage to cloud storage nodes whose read / write performance matches the hotness or coldness of the data objects. This allows flexible use of each cloud storage node based on the hotness or coldness of the data, improving the utilization rate of each cloud storage node. This not only meets user application needs, but also reduces storage costs and effectively improves the overall access performance of the cloud storage system.
[0026]
[0026] The following describes in detail the technical solutions provided by various embodiments of the present disclosure in conjunction with the accompanying drawings.
[0027] FIG1 is a schematic diagram of the structure of a cloud storage system provided by an embodiment of the present disclosure. Referring to FIG1 , the cloud storage system includes a computing layer and a storage layer. The computing layer includes at least one computing node, and the storage layer includes at least two types of cloud storage nodes. The different types of cloud storage nodes have different read and write performances. The read and write performance is related to the data read and write speed of the cloud storage node. Higher read and write performance means faster data can be written to or read from the cloud storage node. Conversely, lower read and write performance means slower data can be written to or read from the cloud storage node.
[0028] In this embodiment, computing nodes include, for example, but are not limited to, read-write nodes, read-only nodes, or write-only nodes. Read-write nodes are computing nodes that can perform both read and write operations, can read and write data in the storage layer, and support operations such as adding, deleting, modifying, or querying. Read-only nodes are computing nodes that can only perform query operations and can only read data from the storage layer but cannot write data to the storage layer. Write-only nodes are computing nodes that can perform write operations and can write data to the storage layer, supporting operations such as adding, deleting, and modifying.
[0029] In this embodiment, the local memory area of the computing node includes, for example, but is not limited to: a local memory buffer pool (BP) and a log buffer. The local memory buffer pool (BP) can cache data pages or index pages in units of pages. The page size is, for example, 16K, 32K, 48K, etc. The data in the data page is usually table data in a database table, and the index page records index information. The log buffer can cache various log pages in units of pages. Log pages include, for example, but are not limited to, log pages recording redo logs (Redo Logs), log pages recording rollback logs (Undo Logs), etc.
[0030] In practical applications, different storage media have different read / write performance and costs. Referring to FIG2 , the order of read / write performance from high to low is: volatile memory, local disk / non-volatile memory, cloud disk, and object-based storage system (OSS). The order of storage cost from high to low is: volatile memory, local disk / non-volatile memory, cloud disk, and object-based storage system. Volatile memory includes, but is not limited to, dynamic random access memory (DRAM) and static random access memory (SRAM). A local disk (also known as a local hard disk) generally refers to a hard disk local to a computing node, including, but not limited to, a solid-state drive (SSD) and a mechanical hard disk. A cloud disk (also known as a cloud hard disk) refers to an online storage service connected to the Internet. Users can store data on the cloud disk and access the data anytime, anywhere through the network. Object storage systems store data in the form of objects and access them through unique identifiers. It is understandable that computing nodes do not need to access local disks through the network. Therefore, the access speed of local disks is relatively faster, that is, the read and write performance is relatively good. Computing nodes usually access data through RDMA.
[0031] Cloud disks can be accessed over a Remote Direct Memory Access (RDMA) network, while compute nodes typically access object storage systems over a TCP (Transmission Control Protocol) network. Compute nodes access cloud disks faster than object storage systems, meaning that cloud disks offer better read and write performance than object storage systems.
[0032] In practical applications, cloud storage can achieve flexible storage capacity compared to local storage. Referring to FIG3 , the storage cost is ranked from high to low as follows: dynamic random access memory, non-volatile memory, solid-state drive, cloud disk, and object storage system. The read and write performance is ranked from high to low as follows: dynamic random access memory, non-volatile memory, solid-state drive, cloud disk, and object storage system. The storage capacity is ranked from small to large as follows: dynamic random access memory, non-volatile memory, solid-state drive, cloud disk, and object storage system. The storage capacity of cloud disks and object storage systems that provide cloud storage services supports flexible expansion.
[0033] In this embodiment, the storage layer is used to persist data. The storage layer includes at least two types of cloud storage nodes, and the different types of cloud storage nodes have different read and write performances. Factors affecting the read and write performance of cloud storage nodes include, but are not limited to, the storage medium and storage technology used by the cloud storage nodes. Cloud storage nodes with different read and write performances use at least one different storage medium and storage technology. For example, in the case where the cloud storage nodes include a cloud disk and an object storage system, the storage medium used by the cloud disk and the storage medium used by the object storage system can be different. In addition, the cloud disk can use a distributed file system for data storage, and the object storage system can use object storage technology for data storage.
[0034] During the phase of writing data from a computing node to a cloud storage node, the computing node determines a first data object and its attribute information that needs to be persistently stored in the cloud storage system, the attribute information including the hotness or coldness of the first data object; selects a first cloud storage node from at least two types of cloud storage nodes whose read / write performance matches the hotness or coldness of the first data object; and writes the first data object to the first cloud storage node. In this embodiment, from a data type perspective, the first data object includes, but is not limited to, table data, index data, redo logs, and rollback logs in a database table. From a page type perspective, the first data object includes, but is not limited to, data pages that record table data, redo log pages that record redo logs, rollback log pages that record rollback logs, and index pages that record index data. Data pages in the local memory cache pool include, but are not limited to, dirty data pages. Dirty data pages are data pages whose content is inconsistent with that of data pages on the cloud storage node. Typically, data pages in the local memory cache pool are read from cloud storage nodes. When a data page in the local memory cache pool is modified, the modified data page has not yet been synchronized to the cloud storage node. At this time, the modified data page in the local memory cache pool is a dirty data page.
[0035] In this embodiment, the attribute information of the first data object includes, but is not limited to, the data type, page type, and popularity of the first data object. The popularity of the first data object can be distinguished by access frequency, data importance, or data size, but is not limited thereto.
[0036] In practical applications, the attribute information of data objects adapted to cloud storage nodes with different read and write performances can be flexibly set as needed. For example, log pages can be adapted to cloud disks, data pages with high access frequency can be adapted to cloud disks, and data pages with low access frequency can be adapted to object storage systems.
[0037] Typically, data objects requiring persistent storage in a computing node are located in the computing node's local memory area. For example, the computing node may receive a first data object from a client and temporarily store it in the local memory area. Then, based on a persistence policy, the computing node persistently stores the first data object in a cloud storage node at an appropriate time or under appropriate conditions. The computing node may persist the data object requiring persistent storage in a cloud storage node whose read / write performance matches the data object's hot / cold level. This allows flexible use of each cloud storage node based on data hot / cold levels, improving the utilization rate of each cloud storage node. This not only meets user application needs, but also reduces storage costs and effectively improves the overall access performance of the cloud storage system. Optionally, the attribute information of the first data object also includes the data type of the first data object; when the computing node selects the first cloud storage node whose read / write performance is adapted to the hot / cold degree of the first data object, it is specifically used to: when the first data object is a redo log, select the cloud storage node with the highest read / write performance from at least two types of cloud storage nodes as the first cloud storage node adapted to the hot / cold degree of the redo log; when the first data object is a data page, select the cloud storage node with the read / write performance adapted to the hot / cold degree of the first data object from at least two types of cloud storage nodes as the first cloud storage node.
[0038] The primary function of the redo log is data recovery. When writing data to a cloud storage system, the data is generally written to the local memory area of a compute node and waits for an appropriate time or condition to be persistently stored in the cloud storage node. Simultaneously, a redo log for the data is generated and promptly written to the cloud storage node. Therefore, the redo log is generally highly popular. To improve data recovery efficiency, the cloud storage node with the highest read / write performance is selected to write the redo log. This allows the redo log to be quickly read from the cloud storage node during the data recovery phase. Of course, other cloud storage nodes can also be selected for writing the redo log as needed. It is understood that writing data pages from a compute node to a cloud storage node with read / write performance that matches the hot / cold nature of the data pages increases the utilization rate of each cloud storage node and reduces storage costs. Preferably, data pages in the local memory area are written to the cloud storage node with the highest read / write performance.
[0039] In practical applications, when writing a first data object to a first cloud storage node, the computing node is specifically configured to: determine a first data page corresponding to the first data object; obtain first metadata information corresponding to the first data page, where the first metadata information includes a storage identifier pointing to the storage node where the first data page is located and a storage path of the first data page on the storage node where the first data page is located; and, if the storage identifier in the first metadata information points to the first cloud storage node, write the first data object to the first data page on the first cloud storage node according to the storage path in the first metadata information. It is understood that the first data page may be a data page or a log page on the storage node.
[0040]
[0039] In practical applications, there are no restrictions on determining the first data page corresponding to the first data object. For example, a mapping relationship between the object identifier of the data object in the local memory area and the page number of the page on the storage node is maintained. The mapping relationship is queried to determine the first data page corresponding to the first data object. For another example, a mapping relationship between the storage address of the data object in the local memory area and the storage address of the page on the storage node is maintained. The mapping relationship is queried to determine the first data page corresponding to the first data object. For another example, the first data page corresponding to the first data object is determined based on an index tree of the cloud storage system. The index tree of the cloud storage system can record index information of data pages in the entire cloud storage system, and the leaf nodes of the index tree record relevant information of the data pages.
[0041]
[0040] Optionally, the computing node is also used to: when the storage identifier in the first metadata information points to the second cloud storage node, write the first data object into an empty data page in the first cloud storage node to obtain a new first data page; update the first metadata information to obtain metadata information corresponding to the new first data page, and the updated first metadata information includes the storage identifier pointing to the first cloud storage node and the storage path of the new first data page on the first cloud storage node.
[0042]
[0041] In actual applications, the first data page may have been transferred from the first cloud storage node to the second cloud storage node. In this case, the first data object may be written to an empty data page in the first cloud storage node to obtain a new first data page.
[0043]
[0042] After the new first data page is written to the first cloud storage node, the first data page on the second cloud storage node will become garbage data over time, and data can be cleared later.
[0044] In some optional embodiments, to improve data security, the computing node is further configured to: upon determining the first data object, add a write lock to the first data object; and upon writing the first data object to the first data page, release the write lock. The purpose of adding the write lock (also known as a mutex lock) is to prevent other programs or processes from modifying the first data object while the first data object is being written to the cloud storage node.
[0045]
[0044] Optionally, metadata information for each data page in at least two types of cloud storage nodes is stored on a target cloud storage node whose read / write performance meets set conditions. When the computing node obtains the first metadata information corresponding to the first data page, it is specifically configured to: obtain the first metadata information corresponding to the first data page locally from the first cloud storage node if the first cloud storage node is the target cloud storage node; and obtain the first metadata information corresponding to the first data page from the target cloud storage node if the first cloud storage node is not the target cloud storage node. Specifically, the set conditions are set as needed and can be set with the goal of efficiently obtaining metadata for each data page. For example, the cloud storage node with the highest read / write performance can be used as the target cloud storage node. For another example, the cloud storage node with the largest storage capacity can be used as the target cloud storage node. For another example, a cloud storage node accessible via an RDMA network can be selected as the target cloud storage node.
[0046]
[0045] In embodiments of the present disclosure, data object migration between cloud storage nodes can also be supported. By transferring data objects from cloud storage nodes with high read / write performance to cloud storage nodes with low read / write performance, the utilization rate of each cloud storage node can be improved and storage costs can be reduced. Based on this, the computing node is further configured to: if a second cloud storage node with lower read / write performance than the first cloud storage node exists among at least two types of cloud storage nodes, transfer a second data object that meets the first transfer condition from the first cloud storage node to the second cloud storage node.
[0047]
[0046] Specifically, the system supports the migration of data objects between cloud storage nodes, thereby improving the utilization rate of each cloud storage node and reducing storage costs. In practical applications, there may be one or more cloud storage nodes with lower read / write performance than the first cloud storage node. If there are multiple cloud storage nodes with lower read / write performance than the first cloud storage node among the at least two types of cloud storage nodes, the computing node is further configured to: select a second cloud storage node from the multiple cloud storage nodes with lower read / write performance than the first cloud storage node. For example, a cloud storage node may be randomly selected from the multiple cloud storage nodes with lower read / write performance than the first cloud storage node as the second cloud storage node; alternatively, the cloud storage node with the highest read / write performance may be selected as the second cloud storage node; or alternatively, the cloud storage node with the lowest read / write performance may be selected as the second cloud storage node, without limitation.
[0048]
[0047] In this embodiment, the first transfer condition can be set as needed. The first transfer condition may be, for example, that the hotness or coldness of the data object does not match the read / write performance of the cloud storage node. For another example, the first transfer condition may be that when the data type of the data object is table data, the hotness or coldness of the data object does not match the read / write performance of the cloud storage node. For another example, the first transfer condition may be that when the page type of the data object is a data page, the hotness or coldness of the data object does not match the read / write performance of the cloud storage node.
[0049]
[0048] Exemplarily, when the computing node transfers the second data object that meets the first transfer condition in the first cloud storage node to the second cloud storage node, it is specifically used to: select the second data object from the first cloud storage node according to the first transfer condition, and the second metadata information corresponding to the second data page to which the second data object belongs includes a storage identifier pointing to the first cloud storage node and a storage path of the second data page on the first cloud storage node; write the second data object to the second cloud storage node, and update the second metadata information, and the updated second metadata information includes a storage identifier pointing to the second cloud storage node and a storage path of the second data page on the second cloud storage node.
[0050] In practical applications, data can be transferred at a page granularity or at a coarser granularity than a page granularity. For example, an object consisting of multiple pages is referred to as a data block, i.e., a data block includes multiple pages. When transferring data at a data block granularity, each data page in the data block corresponding to the second data object is transferred from the first cloud storage node to the second cloud storage node, and the metadata of the transferred data page is modified. This allows subsequent correct access to the data page on the cloud storage node, thus achieving the flexibility of the cloud storage system to support multiple cloud storage nodes of different types.
[0051] For example, a read / write node flushes dirty pages from a memory cache pool to a cloud disk. The cloud disk includes multiple data blocks, each of which includes multiple data pages. Access heat is detected for the data blocks in the cloud disk. If a cold data block is detected, the cold data block in the cloud disk is punched to free up storage space on the cloud disk. The cold data block is then migrated to an object storage system, and the metadata of the data pages in the cold data block is updated. The updated metadata includes a storage identifier pointing to the object storage system and the storage path of the data page in the object storage system. This allows subsequent computing nodes to access the data pages in the cold data block from the object storage system. Data blocks with access heat less than a preset access heat threshold are considered cold data blocks, while data blocks with access heat greater than the preset access heat threshold are considered hot data blocks.
[0051] Optionally, when the computing node transfers the second data object that meets the first transfer condition in the first cloud storage node to the second cloud storage node, the computing node selects the second data object from the first cloud storage node according to the first transfer condition, and the second metadata information corresponding to the second data page to which the second data object belongs includes a storage identifier pointing to the first cloud storage node and a storage path of the second data page on the first cloud storage node; the second data object is written to the local memory area of the computing node, and the second data object in the local memory area is written to the second cloud storage node, and the second metadata information is updated, and the updated second metadata information includes a storage identifier pointing to the second cloud storage node and a storage path of the second data page on the second cloud storage node.
[0052]
[0052] Optionally, the first transfer condition includes a data type screening condition and a hot / cold degree screening condition; when the computing node selects the second data object from the first cloud storage node according to the first transfer condition, it is specifically used to: select, according to the data type screening condition, from the data objects stored in the first cloud storage node, a candidate data object that meets the data type screening condition; and select, according to the hot / cold degree screening condition, from the candidate data objects a data object whose hot / cold degree meets the hot / cold degree screening condition as the second data object.
[0053] Specifically, the data type screening condition includes condition information for selecting data objects, that is, limiting the data types of the data objects to be transferred. For example, the data type screening condition indicates that data objects of the table data type are to be transferred. The hotness screening condition refers to condition information that limits the hotness level that the data objects to be transferred must meet. For example, the hotness screening condition indicates that data objects whose hotness level is less than a preset hotness threshold are to be transferred.
[0054]
[0054] In some optional embodiments, to meet diverse transfer requirements, the computing node is further configured to: monitor whether a trigger event for performing a data transfer operation on the first cloud storage node occurs; if so, trigger an operation to select a second data object from the first cloud storage node based on the first transfer condition. In actual applications, the trigger event is set as needed. Examples of the trigger event include, but are not limited to: the remaining storage capacity of the first cloud storage node being less than a set capacity threshold, the expiration of a set data transfer period, and / or the receipt of a data transfer instruction.
[0055]
[0055] Optionally, the cloud storage system may further include a local storage node, which may be located in a computing node. The local storage node may be, for example, a local disk or a non-volatile memory. The local storage node may serve as a cache system for the cloud storage node, caching data on the cloud storage node, thereby accelerating data access performance of the cloud storage system.
[0056] Based on the above, the computing node is further configured to: write a fourth data object that meets a second transfer condition from the cloud storage node with the highest read / write performance to the local storage node, where the fourth data object corresponds to a fourth data page, and update fourth metadata information corresponding to the fourth data page. The second transfer condition can be flexibly set as needed. For example, the second transfer condition may be to transfer data pages that meet the required hot / cold degree to the cloud storage node with the highest read / write performance. For example, a data page that meets the required hot / cold degree is a data page whose hot / cold degree is greater than a preset hot / cold degree threshold.
[0057] It is understood that after the fourth data page is written to the local storage node, the updated fourth metadata information corresponding to the fourth data page includes a storage identifier pointing to the local storage node where the fourth data page is located and a storage path of the fourth data page on the local storage node where the fourth data page is located. In this way, the fourth data page can be accessed from the local storage node based on the fourth metadata information.
[0058]
[0058] The following describes the stage in which the computing node reads data from the storage node.
[0059] During the data reading phase, the computing node is used to: determine, in response to a received read request, a third data page corresponding to the third data object; obtain third metadata information corresponding to the third data page, where the third metadata information includes a storage identifier pointing to a storage node where the third data page is located and a storage path of the third data page on the storage node where the third data page is located; read the third data object from the storage node pointed to by the storage identifier in the third metadata information according to the storage path in the third metadata information, and provide the third data object to the initiator of the read request.
[0059]
[0060] Optionally, in response to a received read request, the computing node preferentially queries a local memory area of the computing node. If the third data object does not exist in the local memory area, the computing node performs the step of determining a third data page corresponding to the third data object and subsequent steps. If the third data object exists in the local memory area, the computing node reads the third data object from the local memory area and provides the third data object to the sender of the read request, thereby accelerating data reading.
[0060]
[0061] In practical applications, there are no restrictions on determining the third data page corresponding to the third data object. For example, a mapping relationship between the object identifier of the data object in the local memory area and the page number of the page on the storage node is maintained. This mapping relationship is then queried to determine the third data page corresponding to the third data object. Another example is a mapping relationship between the storage address of the data object in the local memory area and the storage address of the page on the storage node is maintained. This mapping relationship is then queried to determine the third data page corresponding to the third data object. Another example is determining the third data page corresponding to the third data object based on an index tree of the cloud storage system. The index tree of the cloud storage system may record index information for data pages across the entire cloud storage system, with leaf nodes of the index tree recording relevant information about the data pages.
[0061]
[0062] In actual applications, the third data page requested to be read in the read request received by the computing node may be in a local storage node, a cloud storage node, or a local memory area. If it is not in the local memory area, the storage node and storage path where the third data page is located are determined based on the metadata information of the third data page, and the local storage node or cloud storage node is accessed according to the storage path of the third data page to obtain the third data page, and the read third data page is returned to the requester who initiated the read request.
[0062]
[0063] The cloud storage system provided by the embodiments of the present disclosure includes at least one computing node and multiple cloud storage nodes with varying read / write performance. The computing node can persist data objects requiring persistent storage to cloud storage nodes with read / write performance matching the data object's hot / cold status. This allows flexible use of each cloud storage node based on data hot / cold status, improving the utilization rate of each cloud storage node. This not only meets user application needs, but also reduces storage costs and effectively improves the overall access performance of the cloud storage system.
[0063]
[0064] To better understand the technical solutions provided by the embodiments of the present disclosure, the following description uses a cloud storage system as an example, using a cloud-native database. Referring to Figure 5 , the cloud-native database includes a computing layer and a storage layer. The computing layer includes read-write nodes and multiple read-only nodes. The storage layer includes a cloud disk and an object storage system. The read-write nodes and read-only nodes include a memory cache pool, a log buffer, and a local disk. The read-write nodes and read-only nodes access the cloud disk via an RDMA network and access the object storage system via a TCP network. Optionally, the cloud-native database may also include a proxy layer, which is responsible for data exchange between clients and computing nodes.
[0064]
[0065] First, let's describe the process of writing data from the compute layer to the storage layer, also known as the data persistence phase. This phase can be performed by read / write nodes, for example, by the storage engine within those nodes. Specifically, the data persistence process includes the following steps:
[0065]
[0066] S 1. The read-write node flushes the first memory data page in the local memory cache pool to the cloud disk.
[0066]
[0067] The first memory data page is a data page in the local memory cache pool that needs to be written to the cloud disk. The first memory data page includes but is not limited to a dirty data page and a data page that has not yet been written to the cloud disk.
[0067]
[0068] A table file in a cloud disk consists of multiple data blocks, each of which contains multiple data pages. Table files, also known as ibd files (table data files), are used to persistently store table data in database tables.
[0068]
[0069] S2. Check whether the cloud disk meets the preset conditions.
[0069]
[0070] S3. If the preset conditions are met, a hole punching process is performed on a first data block among the multiple data blocks included in the cloud disk, and page data in multiple remote data pages under the first data block is migrated to the object storage system.
[0070]
[0071] Optionally, flushing a first memory data page in a local memory cache pool to a cloud disk includes: obtaining metadata information of a first remote data page corresponding to the first memory data page; if the storage medium identifier in the metadata information of the first remote data page is a cloud disk, refreshing the first remote data page using page data in the first memory data page; if the storage medium identifier in the metadata information of the first remote data page is an object storage system, writing the page data in the first memory data page to an empty data page in the cloud disk to obtain a new first remote data page corresponding to the first memory data page; and saving the metadata information of the new first remote data page, wherein the storage medium identifier in the metadata information of the new first cloud disk data page is a cloud disk identifier, and the storage path in the metadata information of the new first cloud disk data page is a storage path for accessing the new first cloud disk data page from the cloud disk. A data page in a cloud disk or object storage system is referred to as a remote data page, and a remote data page is also a data page in the cloud disk or object storage system that records table data. The first remote data page refers to the remote data page corresponding to the first memory data page.
[0071]
[0072] Optionally, migrating page data in multiple remote data pages under the first data block to the object storage system includes: for any first remote data page among the multiple remote data pages under the first data block, reading the first remote data page from the cloud disk to the memory cache pool; writing the first remote data page from the memory cache pool to the object storage system; updating the storage medium identifier in the metadata information of the first remote data page to the object storage system identifier, and updating the storage path in the metadata information of the first remote data page to the storage path for accessing the first remote data page from the object storage system.
[0072]
[0073] Optionally, detecting whether the cloud disk meets the preset condition includes: detecting whether the remaining storage capacity of the cloud disk is less than a preset storage capacity; if so, determining that the cloud disk meets the preset condition; if not, determining that the cloud disk does not meet the preset condition. And / or, performing access popularity detection on data blocks in the cloud disk. If a first data block with an access popularity less than a preset access popularity is detected, determining that the cloud disk meets the preset condition. If not, determining that the cloud disk does not meet the preset condition.
[0073]
[0074] Optionally, the local disk can also serve as a cache system for the cloud disk. Furthermore, the read / write node caches the second remote data page in the cloud disk that meets the cache conditions in the local disk, updates the storage medium identifier in the metadata of the second remote data page cached in the local disk to the local disk identifier, and updates the storage path in the metadata of the second remote data page cached in the local disk to the storage path for accessing the corresponding second remote data from the local disk.
[0074]
[0075] For example, referring to Figure 4 , dirty data pages in the memory cache pool are written to the cloud disk. The cloud disk then performs access popularity detection on the data blocks to identify cold data blocks on the cloud disk. For example, if data block 2 in table file 1 on the cloud disk is a cold data block, hole punching is performed on data block 2 in table file 1 to create a hole file, freeing up storage space on the cloud disk. Data block 2 is then migrated from the cloud disk to the object storage system. In this case, the storage medium identifier in the metadata of each data page in data block 2 is the object storage system identifier, and the storage path in the metadata of each data page in data block 2 is the storage path used to access the data page from the object storage system. For another example, if data block 3 in table file 2 on the cloud disk is a cold data block, hole punching is performed on data block 3 to free up storage space on the cloud disk and migrate data block 3 from the cloud disk to the object storage system. In this case, the storage medium identifier in the metadata of each data page in data block 3 is the object storage system identifier, and the storage path in the metadata of each data page in data block 3 is the storage path for accessing the data page from the object storage system. The local disk on the compute node can also cache some data pages from the cloud disk. The storage medium identifier in the metadata of the data page cached on the local disk is the local disk identifier, and the storage path in the metadata of the data page cached on the local disk is the storage path for accessing the data page from the local disk. The storage medium identifier in the metadata of the data page stored on the cloud disk is the cloud disk identifier, and the storage path in the metadata of the data page stored on the cloud disk is the storage path for accessing the data page from the cloud disk.
[0075]
[0076] Next, we'll describe the stage where a compute node responds to a client's read request and reads data. This stage can be performed by either a read-write node or a read-only node, for example, by the storage engine within a read-write node or a read-only node. Specifically, the data reading process can include the following steps.
[0076]
[0077] The SK computing node responds to the read request sent by the client. If the data to be read does not exist in the memory cache pool, it obtains the metadata information of the target remote data page corresponding to the data to be read.
[0077]
[0078] S2. Write the target remote data page into the memory cache pool according to the metadata information of the target remote data page.
[0078]
[0079] S3. Read the data to be read from the target remote data page in the memory cache pool and return the data to be read to the client.
[0079]
[0080] The computing node may receive a read request from the client through the proxy layer, or the computing node may return the data to be read to the client through the proxy layer.
[0080]
[0081] Optionally, writing the target remote data page to the memory cache pool based on the metadata information of the target remote data page includes: if the storage medium identifier in the metadata information of the target remote data page is a cloud disk identifier, writing the target remote data page from the cloud disk to the memory cache pool based on the storage path in the metadata information of the target remote data page. If the storage medium identifier in the metadata information of the target remote data page is an object storage system identifier, writing the target remote data page from the object storage system to the memory cache pool based on the storage path in the metadata information of the target remote data page. If the storage medium identifier in the metadata information of the target remote data page is a local disk identifier, writing the target remote data page from the local disk to the memory cache pool based on the storage path in the metadata information of the target remote data page.
[0081]
[0082] In some scenarios, cloud disks can relatively ensure fast data read, write, and processing, but storage costs are relatively high. They can better meet the high performance requirements of services. Frequently accessed active data (also known as hot data) can be stored in cloud disks. Less frequently accessed inactive data (also known as cold data) can be stored in object storage systems. While object storage systems offer relatively lower read and write performance, they can significantly reduce storage costs and are suitable for scenarios such as backup, archiving, and long-term storage. Therefore, cloud-native databases utilize a tiered storage strategy for hot and cold data, allowing users to flexibly configure storage resources based on data access frequency and service requirements, meeting performance requirements while minimizing storage costs.
[0082]
[0083] In the cloud-native database provided by the embodiments of the present disclosure, compute nodes can persist data pages requiring persistent storage in the memory cache pool to cloud disks, and data pages in the cloud disks can be migrated to object storage systems. This allows flexible use of cloud disks or object storage systems based on data hotness and coldness, improving their utilization. This not only meets user application needs, but also reduces storage costs and effectively improves the overall access performance of the cloud-native database.
[0083]
[0084] FIG5 is a flowchart of a data storage method provided by an embodiment of the present disclosure. This method is applied to a cloud storage system comprising at least two types of cloud storage nodes, each of which has different read and write performance. Referring to FIG5 , the method may include the following steps.
[0084]
[0085] 501. Determine a first data object and its attribute information that need to be persistently stored in a cloud storage system, where the attribute information includes a hotness or coldness of the first data object.
[0085]
[0086] 502. Select a first cloud storage node whose read / write performance matches the hot / cold degree of the first data object from at least two types of cloud storage nodes, and write the first data object into the first cloud storage node.
[0086]
[0087] Optionally, if there is a second cloud storage node with lower read / write performance than the first cloud storage node among the at least two types of cloud storage nodes, the second data object meeting the first transfer condition in the first cloud storage node is transferred to the second cloud storage node.
[0087]
[0088] Optionally, the attribute information of the first data object also includes the data type of the first data object; selecting a first cloud storage node whose read / write performance is adapted to the hot / cold degree of the first data object from at least two types of cloud storage nodes includes: when the first data object is a redo log, selecting a cloud storage node with the highest read / write performance from at least two types of cloud storage nodes as the first cloud storage node adapted to the hot / cold degree of the redo log; when the first data object is a data page, selecting a cloud storage node whose read / write performance is adapted to the hot / cold degree of the first data object from at least two types of cloud storage nodes as the first cloud storage node.
[0088]
[0089] Optionally, writing the first data object to the first cloud storage node includes: determining the first data page corresponding to the first data object based on the index tree of the cloud storage system; obtaining first metadata information corresponding to the first data page, the first metadata information including a storage identifier pointing to the storage node where the first data page is located and a storage path of the first data page on the storage node where the first data page is located; when the storage identifier in the first metadata information points to the first cloud storage node, writing the first data object to the first data page on the first cloud storage node according to the storage path in the first metadata information.
[0089]
[0090] Optionally, the above method also includes: when the storage identifier in the first metadata information points to the second cloud storage node, writing the first data object into an empty data page in the first cloud storage node to obtain a new first data page; updating the first metadata information to obtain metadata information corresponding to the new first data page, the updated first metadata information including the storage identifier pointing to the first cloud storage node and the storage path of the new first data page on the first cloud storage node.
[0090]
[0091] Optionally, the method further includes: when the first data object is determined, adding a write lock to the first data object; and when the first data object is written to the first data page, releasing the write lock.
[0092] Optionally, metadata information of each data page in at least two types of cloud storage nodes is stored on a target cloud storage node whose read and write performance meets set conditions; obtaining first metadata information corresponding to the first data page includes: when the first cloud storage node is the target cloud storage node, obtaining the first metadata information corresponding to the first data page locally from the first cloud storage node; when the first cloud storage node is not the target cloud storage node, obtaining the first metadata information corresponding to the first data page from the target cloud storage node.
[0091]
[0093] Optionally, selecting a first cloud storage node whose read / write performance is adapted to the attribute information of the first data object from at least two types of cloud storage nodes includes: when there are multiple cloud storage nodes with lower read / write performance than the first cloud storage node among the at least two types of cloud storage nodes, selecting a second cloud storage node from the multiple cloud storage nodes with lower read / write performance than the first cloud storage node; wherein, the method of selecting the second cloud storage node includes: randomly selecting a cloud storage node as the second cloud storage node, or selecting the cloud storage node with the highest read / write performance as the second cloud storage node, or selecting the cloud storage node with the lowest read / write performance as the second cloud storage node.
[0092]
[0094] Optionally, transferring a second data object that meets a first transfer condition in the first cloud storage node to a second cloud storage node includes: selecting the second data object from the first cloud storage node according to the first transfer condition, where the second metadata information corresponding to the second data page to which the second data object belongs includes a storage identifier pointing to the first cloud storage node and a storage path of the second data page on the first cloud storage node; writing the second data object to the second cloud storage node and updating the second metadata information, where the updated second metadata information includes a storage identifier pointing to the second cloud storage node and the storage path of the second data page on the second cloud storage node.
[0093]
[0095] Optionally, the first transfer condition includes a data type screening condition and a hot / cold degree screening condition; and selecting the second data object from the first cloud storage node according to the first transfer condition includes: selecting, from the data objects stored in the first cloud storage node, a candidate data object that meets the data type screening condition according to the data type screening condition; and selecting, from the candidate data objects, a data object whose hot / cold degree meets the hot / cold degree screening condition as the second data object according to the hot / cold degree screening condition.
[0094]
[0096] Optionally, the above method also includes: monitoring whether a trigger event for performing a data transfer operation on the first cloud storage node occurs; the trigger event includes the remaining storage capacity of the first cloud storage node being less than a set capacity threshold, the set data transfer cycle arriving and / or receiving a data transfer instruction; if so, triggering the operation of selecting a second data object from the first cloud storage node according to the first transfer condition.
[0095]
[0097] Optionally, the above method also includes: receiving a read request, the read request being used to request reading a third data object; determining, based on an index tree of the cloud storage system, a third data page corresponding to the third data object; obtaining third metadata information corresponding to the third data page, the third metadata information including a storage identifier pointing to a storage node where the third data page is located and a storage path of the third data page on the storage node where the third data page is located; reading the third data object from the storage node pointed to by the storage identifier in the third metadata information according to the storage path in the third metadata information, and providing the third data object to the initiator of the read request.
[0096]
[0098] Optionally, the above method further includes: if the third data object exists in the local memory area, reading the third data object from the local memory area, and providing the third data object to the sender of the read request.
[0097]
[0099] Optionally, the method further includes: writing a fourth data object that meets the second transfer condition in the cloud storage node with the highest read / write performance to a local storage node in the cloud storage system, the fourth data object corresponding to a fourth data page, and updating fourth metadata information corresponding to the fourth data page.
[0098]
[0100] The implementation manner and technical effects of each step in the above method embodiment have been described in detail in the above system embodiment and will not be elaborated here.
[0099]
[0101] Figure 6 is a flowchart of another data reading method provided by an embodiment of the present disclosure. This method is applied to a cloud storage system comprising at least two types of cloud storage nodes, each of which has different read and write performance. Referring to Figure 6 , the method may include the following steps:
[0100]
[0102] 60K receives read requests, which are used to request to read data objects.
[0101]
[0103] 602. Determine a data page corresponding to the data object according to the index tree of the cloud storage system.
[0102]
[0104] 603. Obtain metadata information corresponding to the data page, where the metadata information includes a storage identifier pointing to the storage node where the data page is located and a storage path of the data page on the storage node where the data page is located.
[0103]
[0105] 604. Read the data object from the storage node pointed to by the storage identifier in the metadata information according to the storage path in the metadata information, and provide the data object to the initiator of the read request.
[0104]
[0106] The storage node where the data page is located is any type of cloud storage node or a local storage node of the cloud storage system.
[0105]
[0107] The implementation manner and technical effects of each step in the above method embodiment have been described in detail in the above system embodiment and will not be elaborated here.
[0106]
[0108] FIG7 is a schematic diagram of the structure of a data access device provided in an embodiment of the present disclosure. The data access device is applied to a cloud storage system including at least two types of cloud storage nodes, each of which has different read and write performance. Referring to FIG7 , the device may include:
[0107]
[0109] A determination module 71 is configured to determine a first data object and its attribute information that needs to be persistently stored in the cloud storage system, where the attribute information includes a hotness or coldness of the first data object;
[0108]
[0110] The selection module 72 is configured to select a first cloud storage node having read / write performance that is compatible with the hot / cold degree of the first data object from at least two types of cloud storage nodes.
[0109]
[0111] A writing module 73, configured to write the first data object into the first cloud storage node;
[0110]
[0112] Optionally, the above-mentioned device also includes: a transfer module 74, which is used to transfer the second data object that meets the first transfer condition in the first cloud storage node to the second cloud storage node when there is a second cloud storage node with lower read and write performance than the first cloud storage node among at least two types of cloud storage nodes.
[0111]
[0113] Optionally, the attribute information of the first data object also includes the data type of the first data object; the selection module 72 is specifically used to: when the first data object is a redo log, select a cloud storage node with the highest read / write performance from at least two types of cloud storage nodes as the first cloud storage node adapted to the hot / cold degree of the redo log; when the first data object is a data page, select a cloud storage node with read / write performance adapted to the hot / cold degree of the first data object from at least two types of cloud storage nodes as the first cloud storage node.
[0114] Optionally, the writing module 73 is specifically configured to: determine, based on an index tree of the cloud storage system, a first data page corresponding to the first data object; obtain first metadata information corresponding to the first data page, where the first metadata information includes a storage identifier pointing to a storage node where the first data page is located and a storage path of the first data page on the storage node where the first data page is located; and, when the storage identifier in the first metadata information points to the first cloud storage node, write the first data object to the first data page on the first cloud storage node according to the storage path in the first metadata information.
[0112]
[0115] Optionally, the writing module 73 is further configured to: when the storage identifier in the first metadata information points to the second cloud storage node, write the first data object to an empty data page in the first cloud storage node to obtain a new first data page; and update the first metadata information to obtain metadata information corresponding to the new first data page, where the updated first metadata information includes the storage identifier pointing to the first cloud storage node and the storage path of the new first data page on the first cloud storage node.
[0113]
[0116] Optionally, the writing module 73 is further configured to: add a write lock to the first data object when the first data object is determined; and release the write lock when the first data object is written to the first data page.
[0114]
[0117] Optionally, metadata information of each data page in at least two types of cloud storage nodes is stored on a target cloud storage node whose read and write performance meets set conditions; when the write module 73 obtains the first metadata information corresponding to the first data page, it is specifically used to include: when the first cloud storage node is the target cloud storage node, obtaining the first metadata information corresponding to the first data page locally from the first cloud storage node; when the first cloud storage node is not the target cloud storage node, obtaining the first metadata information corresponding to the first data page from the target cloud storage node.
[0115]
[0118] Optionally, the selection module 72 is specifically used to: when there are multiple cloud storage nodes with lower read and write performance than the first cloud storage node among at least two types of cloud storage nodes, select a second cloud storage node from the multiple cloud storage nodes with lower read and write performance than the first cloud storage node; wherein the method of selecting the second cloud storage node includes: randomly selecting a cloud storage node as the second cloud storage node, or selecting the cloud storage node with the highest read and write performance as the second cloud storage node, or selecting the cloud storage node with the lowest read and write performance as the second cloud storage node.
[0116]
[0119] Optionally, the transfer module 74 is specifically configured to: select a second data object from the first cloud storage node according to the first transfer condition, where the second metadata information corresponding to the second data page to which the second data object belongs includes a storage identifier pointing to the first cloud storage node and a storage path of the second data page on the first cloud storage node; write the second data object to the second cloud storage node, and update the second metadata information, where the updated second metadata information includes a storage identifier pointing to the second cloud storage node and the storage path of the second data page on the second cloud storage node.
[0117]
[0120] Optionally, the first transfer condition includes a data type screening condition and a hot / cold degree screening condition; when the transfer module 74 selects the second data object from the first cloud storage node according to the first transfer condition, it is specifically configured to: select, from the data objects stored in the first cloud storage node, a candidate data object that meets the data type screening condition according to the data type screening condition; and select, from the candidate data objects, a data object whose hot / cold degree meets the hot / cold degree screening condition as the second data object according to the hot / cold degree screening condition.
[0118]
[0121] Optionally, the transfer module 74 is further used to: monitor whether a trigger event for performing a data transfer operation on the first cloud storage node occurs; the trigger event includes that the remaining storage capacity of the first cloud storage node is less than a set capacity threshold, the set data transfer cycle is reached and / or a data transfer instruction is received; if so, trigger the operation of selecting a second data object from the first cloud storage node according to the first transfer condition.
[0119]
[0122] Optionally, the apparatus further includes: a receiving module, configured to receive a read request, where the read request is used to request to read a third data object;
[0120]
[0123] The reading module is configured to determine, based on an index tree of the cloud storage system, a third data page corresponding to the third data object; obtain third metadata information corresponding to the third data page, the third metadata information including a storage identifier pointing to a storage node where the third data page is located and a storage path of the third data page on the storage node where the third data page is located; read the third data object from the storage node pointed to by the storage identifier in the third metadata information according to the storage path in the third metadata information, and provide the third data object to an initiator of the read request.
[0121]
[0124] Optionally, the reading module is further configured to, if the third data object exists in the local memory area, read the third data object from the local memory area and provide the third data object to the sender of the read request.
[0122]
[0125] Optionally, the writing module 73 is further configured to write a fourth data object that meets the second transfer condition in the cloud storage node with the highest read / write performance to a local storage node in the cloud storage system, where the fourth data object corresponds to a fourth data page, and update fourth metadata information corresponding to the fourth data page.
[0123]
[0126] The implementation method and technical effects of each module in the above device embodiment have been described in detail in the above system embodiment and will not be elaborated here.
[0124]
[0127] It should be noted that the execution entity of each step of the method provided in the above embodiment may be the same device, or the method may be executed by different devices. For example, the execution entity of steps 601 to 604 may be device A; for another example, the execution entity of steps 601 and 602 may be device A, and the execution entity of steps 603 and 604 may be device B; and so on.
[0125]
[0128] Furthermore, some of the processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may be executed in a different order than the order in which they appear herein or in parallel. Operation sequence numbers, such as 601 and 602, are merely used to distinguish between different operations and do not represent any specific execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that terms such as "first" and "second" are used herein to distinguish between different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0126]
[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0127]
[0130] FIG8 is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. As shown in FIG8, the computer device includes: a memory 81 and a processor 82;
[0128]
[0131] The memory 81 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, images, videos, etc.
[0129]
[0132] The memory 81 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0130]
[0133] The processor 82 is coupled to the memory 81, and is configured to execute the computer program in the memory 81, so as to execute the steps in the data storage method or the steps in the data reading method.
[0131]
[0134] Furthermore, as shown in FIG8 , the computer device also includes other components, such as a communication component 83, a display 84, a power supply component 85, and an audio component 86. FIG8 only schematically illustrates some components, and does not mean that the computer device only includes the components shown in FIG8 . Furthermore, the components within the dashed boxes in FIG8 are optional, not mandatory, components, and their specific requirements depend on the product form factor of the computer device. The computer device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smartphone, or an IoT (Internet of Things) device, or as a server-side device such as a conventional server, a cloud server, or a server array. If the computer device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smartphone, it may include the components within the dashed boxes in FIG8 ; if the computer device of this embodiment is implemented as a server-side device such as a conventional server, a cloud server, or a server array, it may not include the components within the dashed boxes in FIG8 .
[0132]
[0135] The detailed implementation process of the processor executing each action can be found in the related description in the aforementioned method embodiment or device embodiment, which will not be repeated here.
[0133]
[0136] Accordingly, an embodiment of the present disclosure further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps in the above method embodiments that can be executed by a computer device.
[0134]
[0137] Accordingly, an embodiment of the present disclosure further provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in the above method embodiment that can be performed by a computer device.
[0135]
[0138] The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2nd Generation), 3G (3rd Generation), 4G (4th Generation) / LTE (Long Term Evolution), 5G (5th Generation), and other mobile communication networks, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0139] The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide action, but also the duration and pressure associated with the touch or slide action.
[0136]
[0140] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.
[0137]
[0141] The audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals may be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0138]
[0142] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139]
[0143] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process flow and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing device, produce means for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams.
[0140]
[0144] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0141]
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0142]
[0146] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.
[0147] Memory may include non-permanent storage in a computer-readable medium, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0143]
[0148] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0144]
[0149] It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or apparatus. In the absence of further limitations, the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus comprising the elements.
[0145]
[0150] The above are merely examples of the present disclosure and are not intended to limit the present disclosure. Persons skilled in the art will readily appreciate that various modifications and variations are possible with the present disclosure. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure are intended to be encompassed by the claims of the present disclosure.
Claims
Claims 1. A data storage method, applied to a cloud storage system, wherein the cloud storage system includes at least two types of cloud storage nodes, and the different types of cloud storage nodes have different read and write performances. The method comprises: Determining a first data object and attribute information thereof that needs to be persistently stored in the cloud storage system, wherein the attribute information includes a hotness or coldness of the first data object; A first cloud storage node having read / write performance that matches the hotness or coldness of the first data object is selected from the at least two types of cloud storage nodes, and the first data object is written into the first cloud storage node.
2. The method according to claim 1, further comprising: In the event that there is a second cloud storage node with lower read and write performance than the first cloud storage node among the at least two types of cloud storage nodes, the second data object that meets the first transfer condition in the first cloud storage node is transferred to the second cloud storage node.
3. The method according to claim 1, wherein: The attribute information of the first data object also includes the data type of the first data object; selecting a first cloud storage node whose read / write performance is adapted to the hot / cold degree of the first data object from the at least two types of cloud storage nodes, including: in a case where the first data object is a redo log, selecting a cloud storage node with the highest read / write performance from the at least two types of cloud storage nodes as the first cloud storage node adapted to the hot / cold degree of the redo log; in a case where the first data object is a data page, selecting a cloud storage node whose read / write performance is adapted to the hot / cold degree of the first data object from the at least two types of cloud storage nodes as the first cloud storage node.
4. The method according to claim 1, wherein: Writing the first data object into the first cloud storage node includes: determining the first data page corresponding to the first data object according to the index tree of the cloud storage system; obtaining first metadata information corresponding to the first data page, the first metadata information including a storage identifier pointing to the storage node where the first data page is located and a storage path of the first data page on the storage node where the first data page is located; when the storage identifier in the first metadata information points to the first cloud storage node, writing the first data object into the first data page on the first cloud storage node according to the storage path in the first metadata information.
5. The method according to claim 4, further comprising: When the storage identifier in the first metadata information points to the second cloud storage node, the first data object is written into an empty data page in the first cloud storage node to obtain a new first data page; the first metadata information is updated to obtain metadata information corresponding to the new first data page, and the updated first metadata information includes the storage identifier pointing to the first cloud storage node and the storage path of the new first data page on the first cloud storage node.
6. The method according to claim 4, wherein: The metadata information of each data page in the at least two types of cloud storage nodes is stored on a target cloud storage node whose read and write performance meets the set conditions; Acquiring first metadata information corresponding to the first data page includes: acquiring the first metadata information corresponding to the first data page locally from the first cloud storage node when the first cloud storage node is the target cloud storage node; In a case where the first cloud storage node is not the target cloud storage node, first metadata information corresponding to the first data page is obtained from the target cloud storage node.
7. The method according to claim 2, further comprising: In the case where there are multiple cloud storage nodes with lower read and write performance than the first cloud storage node among the at least two types of cloud storage nodes, selecting a second cloud storage node from the multiple cloud storage nodes with lower read and write performance than the first cloud storage node; The method of selecting the second cloud storage node includes: randomly selecting a cloud storage node as the second cloud storage node, or selecting the cloud storage node with the highest read / write performance as the second cloud storage node, or selecting the cloud storage node with the lowest read / write performance as the second cloud storage node.
8. The method according to claim 2, wherein: Transferring a second data object that meets a first transfer condition in the first cloud storage node to the second cloud storage node includes: selecting the second data object from the first cloud storage node according to the first transfer condition, where second metadata information corresponding to a second data page to which the second data object belongs includes a storage identifier pointing to the first cloud storage node and a storage path of the second data page on the first cloud storage node; writing the second data object to the second cloud storage node and updating the second metadata information, where the updated second metadata information includes a storage identifier pointing to the second cloud storage node and a storage path of the second data page on the second cloud storage node.
9. The method according to any one of claims 1 to 8, further comprising: A read request is received, the read request being used to request reading a third data object; a third data page corresponding to the third data object is determined based on an index tree of the cloud storage system; third metadata information corresponding to the third data page is obtained, the third metadata information including a storage identifier pointing to a storage node where the third data page is located and a storage path of the third data page on the storage node where the third data page is located; based on the storage path in the third metadata information, the third data object is read from the storage node pointed to by the storage identifier in the third metadata information, and the third data object is provided to the initiator of the read request.
10. The method according to any one of claims 1 to 8, further comprising: The fourth data object that meets the second transfer condition in the cloud storage node with the highest read / write performance is written to the local storage node in the cloud storage system, the fourth data object corresponding to the fourth data page, and the fourth metadata information corresponding to the fourth data page is updated.
11. A data reading method, applied to a cloud storage system, wherein the cloud storage system includes at least two types of cloud storage nodes, and the different types of cloud storage nodes have different read and write performances. The method comprises: receiving a read request, wherein the read request is used to request reading of a data object; According to the index tree of the cloud storage system, determine the data page corresponding to the data object; obtain metadata information corresponding to the data page, the metadata information includes a pointer to the storage where the data page is located; the storage identifier of the node and the storage path of the data page on the storage node where the data page is located; reading the data object from the storage node pointed to by the storage identifier in the metadata information according to the storage path in the metadata information, and providing the data object to the initiator of the read request; The storage node where the data page is located is any type of cloud storage node.
12. A cloud storage system, comprising: At least one computing node and at least two types of cloud storage nodes, where different types of cloud storage nodes have different read and write performances; The computing node is configured to determine a first data object that needs to be persistently stored and attribute information thereof, wherein the attribute information includes a hotness or coldness of the first data object; A first cloud storage node having read / write performance that matches the hotness or coldness of the first data object is selected from the at least two types of cloud storage nodes.
13. The cloud storage system according to claim 12, wherein: The computing node is also used to: when there is a second cloud storage node with lower read and write performance than the first cloud storage node among the at least two types of cloud storage nodes, transfer the second data object that meets the first transfer condition in the first cloud storage node to the second cloud storage node; or, in response to a received read request, determine the third data page corresponding to the third data object according to the index tree of the cloud storage system; obtain third metadata information corresponding to the third data page, the third metadata information including a storage identifier pointing to the storage node where the third data page is located and a storage path of the third data page on the storage node where the third data page is located; read the third data object from the storage node pointed to by the storage identifier in the third metadata information according to the storage path in the third metadata information, and provide the third data object to the initiator of the read request.
14. A computer device comprising: A memory and a processor; wherein the memory is used to store a computer program; the processor is coupled to the memory, and is used to execute the computer program to perform the steps in the method according to any one of claims 1 to 11.
15. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 11.
16. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Application data management method and system and computer equipment
CN113835616A
ORAM data reliable storage method and system in heterogeneous cloud storage environment
CN117555486A
Adaptive querying of time-series data over tiered storage
US11461347B1
Tiered storage optimization and migration
US20200326871A1
Selecting Storage Resources Based On Data Characteristics
US20230244399A1