Data processing method and device

By generating joint snapshots of cloud disk storage devices at the file system level, and using the target associated objects in the object storage service, the performance problems caused by full database backup and restoration of dependent database snapshots are solved, and low-cost consistent joint snapshots are realized, ensuring full backup and restore functions of the database.

CN120316070BActive Publication Date: 2025-09-02ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510782823.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-02
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the prior art, full backup and restore of databases depend on database snapshots, resulting in the impact of database performance. After migrating to object storage services, the native snapshot of cloud disk no longer contains complete data, affecting the full backup and restore function.

Method used

By generating joint snapshots of cloud disk storage devices at the file system level, using the target associated objects in the object storage service, a low-cost consistent joint snapshot is achieved, ensuring the full backup and restore function of the database.

Benefits of technology

While reducing database storage costs, the full backup and restore functions are retained to ensure database performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316070B_ABST
    Figure CN120316070B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a data processing method and apparatus, which are applied to a file system in a database, wherein the data processing method includes: in response to a migration instruction of a first file block in a target file, migrating the first file block from a cloud disk storage device to an object storage service, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block; generating a joint snapshot of the cloud disk storage device, wherein the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to a corresponding target-associated object in the object storage service, and the target-associated object is the first file block stored in the object storage service. The metadata of the first file block in the joint snapshot can directly point to the corresponding target-associated object in the object storage service, thereby realizing a low-cost consistent joint snapshot at the file system level and ensuring database performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of database technology, and more particularly to a data processing method and device. Background Art

[0002] With the rapid development of computer and network technologies, database technology has also rapidly evolved. With its abundant hardware resources, high processing performance, and ability to meet the demands of real-time big data analysis, it has been widely adopted in many industries and fields. Existing technologies rely on database snapshots for full database backup and restore. The integrity and consistency of database snapshots can significantly impact database performance. Therefore, there is an urgent need for a low-cost, high-performance database data processing solution. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0004] According to a first aspect of an embodiment of this specification, a data processing method is provided, which is applied to a file system in a database, comprising:

[0005] In response to a migration instruction for a first file block in a target file, migrate the first file block from a cloud disk storage device to an object storage service, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block;

[0006] Generate a joint snapshot of the cloud disk storage device, wherein the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to a corresponding target-associated object in the object storage service, and the target-associated object is the first file block stored in the object storage service.

[0007] According to a second aspect of an embodiment of this specification, there is provided a data processing device, applied to a file system in a database, comprising:

[0008] a migration module configured to migrate a first file block in a target file from a cloud disk storage device to an object storage service in response to a migration instruction for the first file block, wherein the target file is divided into at least one file block and the first file block is any one of the at least one file block;

[0009] A generation module is configured to generate a joint snapshot of the cloud disk storage device, wherein the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to a corresponding target-associated object in the object storage service, and the target-associated object is the first file block stored in the object storage service.

[0010] According to a third aspect of an embodiment of this specification, a computing device is provided, including:

[0011] memory and processor;

[0012] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.

[0013] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.

[0014] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.

[0015] One embodiment of the present specification provides a data processing method, which is applied to a file system in a database, and in response to a migration instruction of a first file block in a target file, migrates the first file block from a cloud disk storage device to an object storage service, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block; and generates a joint snapshot of the cloud disk storage device, wherein the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to a corresponding target-associated object in the object storage service, and the target-associated object is the first file block stored in the object storage service.

[0016] An embodiment of the present specification implements that a file system in a database can migrate the first file block in a target file from a cloud disk storage device to an object storage service to reduce the storage cost of the database, and can generate a joint snapshot of the cloud disk storage device. The metadata of the first file block in the joint snapshot can directly point to the corresponding target associated object in the object storage service, thereby implementing a low-cost consistent joint snapshot at the file system level. The joint snapshot contains complete file data, and based on the joint snapshot, functions such as full backup and restore of the database can be implemented. While reducing the storage cost of the database, important functions such as full backup and restore of the database can be retained, thereby ensuring database performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of a data processing method provided by one embodiment of this specification;

[0018] Figure 2 This is a schematic diagram of a joint snapshot of a cloud disk storage device provided by an embodiment of this specification;

[0019] Figure 3 This is a schematic diagram of a joint snapshot of another cloud disk storage device provided by an embodiment of this specification;

[0020] Figure 4 This is a schematic diagram of a joint snapshot in a database data processing method provided by one embodiment of this specification;

[0021] Figure 5 This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification;

[0022] Figure 6 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0023] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0024] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0025] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0026] First, the terms involved in one or more embodiments of this specification are explained.

[0027] Database: Based on shared storage, it features high availability, elastic scalability, high performance, security and reliability. It can be widely used in e-commerce, finance, government affairs and other fields, helping enterprises cope with the challenges of massive data storage and high concurrent access.

[0028] Shared storage virtual block devices: Similar to cloud drives, they allow the same virtual device to be mounted on different physical machines, enabling shared access to the same storage space. Snapshots can be taken at any time while reading or writing data in a shared storage virtual block device. Each shared storage virtual block device snapshot is a separate, read-only virtual block device, and the data on it reflects the state at the time the snapshot was taken. A read-write shared storage virtual block device is typically called a master-version shared storage virtual block device, and a snapshot generated from a master-version shared storage virtual block device at a specific moment is called a snapshot shared storage virtual block device.

[0029] OSS (Object Storage Service) is a massive, secure, low-cost, and highly reliable cloud storage service. It is a distributed cloud storage solution provided by cloud vendors, significantly lower in cost than block storage, cloud disks, and shared storage virtual block devices. OSS core components include buckets and objects. Buckets are containers for data storage and are globally uniquely named (e.g., "my-website-bucket"). When creating a bucket, you must select a region and storage type. Objects are the basic unit of storage and consist of a key (a unique identifier for the object, such as "images / logo.png"); data (the object's content, arbitrary binary data); and metadata (the object's attributes, such as "Content-Type" and custom metadata).

[0030] File system: A file system that implements file semantics on the storage space of shared storage virtual block devices within the database.

[0031] It should be noted that more and more databases are trying to migrate cold data on cloud disks to OSS to reduce costs. However, after some data is migrated to OSS, the native snapshots of the cloud disks no longer contain complete data, which poses challenges to database full backup and restore functions.

[0032] For example, full database backup and restore rely on shared storage virtual block device snapshots provided by the server cluster and related components of the shared storage virtual block device. To reduce costs, the database can use the file system to migrate less frequently accessed table files or less frequently accessed data blocks within table files on the shared storage virtual block device to OSS. In this case, the snapshot provided by the shared storage virtual block device provider no longer contains complete data, posing challenges to database full backup and restore functions.

[0033] During implementation, some file systems do not have their own snapshot functionality and must rely on snapshots of block devices. For example, XFS is a high-performance log file system with scalability, high performance, and powerful log functionality. It can handle large-capacity data storage and high-concurrency read-write scenarios, and is widely used in data centers, servers, and other fields. Some file systems have their own snapshot functionality, but only support data persistence on block devices. For example, ZFS is a file system designed to handle very large amounts of data. It uses checksum technology to calculate a checksum for each data block and metadata. When reading data, the checksum is automatically verified to ensure that the data has not been damaged or tampered with during storage and transmission. Even if hardware failures or software errors occur, data can be detected and repaired in a timely manner. Snapshots can be created for file systems or data sets. Snapshots are read-only copies of the file system state at a certain moment. They are fast to create and do not take up a lot of additional storage space because they only record the data blocks that have changed since the last snapshot. Some file systems have their own snapshot capabilities, but they only support data persistence on OSS. For example, JuiceFS is a high-performance distributed file system designed for cloud native. It adopts a "data" and "metadata" separate storage architecture. File data is segmented and stored in object storage (such as Amazon S3), and metadata can be stored in Redis, MySQL, TiKV, SQLite and other databases. Users can choose according to the scenario and performance requirements.

[0034] Therefore, one embodiment of this specification provides a data processing method that is applied to the file system in the database, which can generate a consistent joint snapshot of a shared storage virtual block device and an OSS object, and realizes low-cost consistent joint snapshots at the file system level, allowing database customers to reduce costs while retaining important functions such as full backup and restore.

[0035] In this specification, a data processing method is provided. This specification also relates to a data processing apparatus, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0036] See also Figure 1 , Figure 1A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which is applied to a file system in a database and specifically includes the following steps 102 to 104 .

[0037] Step 102: In response to a migration instruction for a first file block in a target file, migrate the first file block from a cloud disk storage device to an object storage service, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block.

[0038] It's important to note that a database is a warehouse that organizes, stores, and manages data according to a data structure. In computers, a database refers to an organized, shareable collection of data stored permanently within a computer. Data in a database is organized, described, and stored according to specific mathematical models, with minimal redundancy, high data independence, and easy scalability, and can be shared by a variety of users.

[0039] The file system is a system within the database that implements file semantics on the storage space of cloud disk storage devices. It is responsible for efficient data organization and management on cloud disk storage devices. The file system can provide multiple interfaces (such as data writing interface and data migration interface) to implement data processing of corresponding files on the storage space of cloud disk storage devices.

[0040] The target file is a file stored on the cloud storage device through the database's file system. It can be any type of file, such as documents, images, and videos. The target file is divided into one or more smaller parts, called file blocks. Dividing files into file blocks facilitates more flexible file management and operations, such as storage and transfer.

[0041] The first file block refers to the file block that meets the specific migration criteria among all the divided file blocks. There may be more than one first file block, and any file block that meets the migration criteria can be the first file block. The first file block can be part of or all of the file blocks in the target file. The migration criteria is a judgment standard used to determine which file blocks need to be migrated. For example, the migration criteria can be the cold data in each first file block. Cold data generally refers to data that is infrequently accessed or used, as opposed to "hot data" (data that is frequently accessed and used). For example, some historical archival data in an enterprise may only be accessed once every few months or even years, and this type of data can be considered cold data.

[0042] Cloud storage devices refer to devices that provide cloud storage space, such as block storage, cloud disks, server clusters, and the shared storage virtual block devices provided by related components. Object storage services are a type of cloud storage service that offers significantly lower storage costs than block storage, cloud disks, server clusters, and the shared storage virtual block devices provided by related components.

[0043] In actual implementation, the file system can split the target file into multiple file blocks and store them in a cloud disk storage device. In response to a migration instruction for the first file block in the target file, the file system can migrate the first file block from the cloud disk storage device to the object storage service to reduce database storage costs. The migration instruction for the first file block can be received as an external trigger instruction, or it can be automatically triggered when it is determined that the first file block meets the set migration instruction.

[0044] Specifically, the database can write the target file to the cloud disk storage device through the data writing interface provided by the file system, and the file system can also provide a data migration interface, and the database can migrate the first file block from the cloud disk storage device to the object storage service through the data migration interface provided by the file system.

[0045] It should be noted that the file data and metadata written to the cloud disk storage device by the database through the data writing interface provided by the file system are all stored on the cloud disk storage device. The file system can also provide the database with a data migration interface, allowing part or all of the file data to be migrated from the cloud disk storage device to the object storage service (OSS) at the granularity of blocks (such as 4MB) to reduce costs, or migrated back to the cloud disk storage device when needed.

[0046] Step 104: Generate a joint snapshot of the cloud disk storage device, where the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to the corresponding target associated object in the object storage service, and the target associated object is the first file block stored in the object storage service.

[0047] Specifically, a joint snapshot refers to a consistent snapshot obtained by combining the data snapshot of the cloud disk storage device and the target associated object pointed to by the first file block migrated to the object storage service. The metadata of the first file block in the joint snapshot points to the corresponding target associated object in the object storage service, and the metadata of the second file block stored on the cloud disk storage device indicates that the corresponding data is stored on the cloud disk storage device. The joint snapshot contains complete file data and can realize important functions such as full backup and restoration of the database, thereby reducing database storage costs while retaining important functions such as full backup and restoration.

[0048] It should be noted that the joint snapshot of the cloud disk storage device is mainly used for full backup and restoration of file data. The file metadata of the file system is always stored on the cloud disk storage device. When the data of the first file block in the target file is migrated to the object storage service, the metadata of this first file block on the cloud disk storage device will point to the target-related object on the object storage service. Therefore, the file system needs to save the cloud disk storage device snapshot at each moment, as well as copies of the target-related objects in all object storage services associated at the historical moment of the snapshot. That is, "cloud disk storage device snapshot + target-related object copy associated with the cloud disk storage device snapshot" constitutes a consistent joint snapshot.

[0049] For example, Figure 2 This is a schematic diagram of a joint snapshot of a cloud disk storage device provided by an embodiment of this specification, such as Figure 2 As shown, the joint snapshot of the cloud disk storage device includes metadata of each file block. Each file block is a 4MB data block. File block 1, file block 3, and file block 4 are stored on the cloud disk storage device. The metadata of file block 2 points to the corresponding target associated object "Object 2" in the object storage service (OSS), and the metadata of file block 4 points to the corresponding target associated object "Object 4" in the object storage service (OSS).

[0050] It should be noted that a joint snapshot of the cloud disk storage device can be generated. The metadata of the first file block in the joint snapshot can directly point to the corresponding target associated object in the object storage service, realizing a low-cost consistent joint snapshot at the file system level. Based on the joint snapshot, functions such as full backup and restore of the database can be realized. While reducing the database storage cost, important functions such as full backup and restore of the database can be retained to ensure database performance.

[0051] In an optional implementation of this embodiment, generating a joint snapshot of the cloud disk storage device includes:

[0052] Generate a first joint snapshot, and point the metadata of the first file block in the first joint snapshot to the target associated object in the object storage service. The first joint snapshot is a snapshot generated at any time, and the metadata of the first file block in the joint snapshots at different times may point to the same or different target associated objects.

[0053] It should be noted that if the target associated objects associated with the metadata of each first file block in the cloud disk storage device snapshot can be copied at the moment of generating the cloud disk storage device snapshot, a joint snapshot can be generated. However, generating a cloud disk storage device snapshot is a process, and copying the target associated objects associated with the metadata of each first file block takes time, which may affect performance and is not conducive to cost control.

[0054] In actual implementation, the joint snapshot of the cloud disk storage device essentially needs to ensure the consistency of the file metadata and the corresponding target-related objects in the joint snapshot at the file system level, and the consistency of the file data and the related target-related object data in the joint snapshot of the cloud disk storage device at the database level.

[0055] Therefore, in an embodiment of the present specification, when generating the first joint snapshot, the metadata of the first file block in the first joint snapshot can be pointed to the target associated object in the object storage service. That is, when generating the first joint snapshot, the metadata of the first file block is temporarily pointed to the current latest target associated object in the object storage service. When the write-time copy is subsequently performed, the metadata of the first file block is modified to be associated with the target associated object copy. In this way, the target associated objects pointed to by the metadata of the first file block in the joint snapshots at different times are the same or different.

[0056] It should be noted that snapshot is a core technology for recording the data status at a certain point in time in a computer system. A snapshot is a static read-only copy of file data at a specific point in time, preserving the complete status of the file data at that moment. Even if the data changes subsequently, the snapshot content remains unchanged.

[0057] In specific implementation, the snapshot version number is incremented. New joint snapshots can be created periodically and the oldest joint snapshot can be deleted. The file system can provide interfaces for creating and deleting snapshots. When creating and deleting joint snapshots through these interfaces, the file system will call the interfaces for creating and deleting joint snapshots accordingly. At the same time, all joint snapshot information is persisted to a fixed area in the storage space of the main version joint snapshot.

[0058] For example, a master-version joint snapshot of a shared storage virtual block device (also known as a master-version block device) provided by a server cluster and related components is typically referred to as "PBD-1." Joint snapshots created from "PBD-1" (also known as read-only block devices of different versions) are typically referred to as "PBD-2," "PBD-3," and so on, where 2 and 3 represent the joint snapshot versions. Directly calling the interfaces provided by the server cluster and related components to create and delete joint snapshots is unaware of the file system. However, if the file system provides interfaces for creating and deleting joint snapshots, the file system will internally call the interfaces provided by the server cluster and related components when creating and deleting snapshots. All snapshot information will also be persisted to a fixed area within the storage space of the master-version PBD snapshot.

[0059] It's important to note that the snapshot information persisted by the file system in the master-version joint snapshot, combined with the post-snapshot copy-on-write (pre-modification and pre-deletion) mechanism for the target object associated with the first file block migrated to the Object Storage service, allows the metadata of the first file block in the file system to maintain information about the target object replica corresponding to the first file block in each joint snapshot. If the snapshot process prohibits modification and deletion of objects in the Object Storage service, that is, suspends copy-on-write, overall data consistency can be guaranteed.

[0060] In actual implementation, each first file block in the target file migrated to the Object Storage Service (OSS) may correspond to a separate target-associated object in each joint snapshot, or the first file block may point to the same target-associated object in multiple joint snapshots. This is key metadata that the file system needs to maintain on the joint snapshot.

[0061] For example, Figure 3 This is a schematic diagram of a joint snapshot of another cloud disk storage device provided by an embodiment of this specification, such as Figure 3 As shown, the dotted box on the left represents the master version joint snapshot (that is, the PBD-1 master version data). The master version joint snapshot includes the metadata of the target file recorded inside the file system. Each box in the metadata represents a file block of the file. The file block with a solid border indicates that the latest data of this file block is on the cloud disk storage device. The file block with a dotted border indicates that the latest data of this file block is migrated to the object storage service (OSS), pointing to the corresponding target associated object on the OSS. PBD-2, PBD-3, ..., PBD-7, etc. in the upper right are joint snapshots generated by the file system at different times. As shown Figure 3 As shown in the figure, after blk_x (the first file block) is migrated to the Object Storage Service (OSS), the first joint snapshot (PBD-2) is generated. At this time, the metadata of the first file block blk_x in the first joint snapshot (PBD-2) temporarily points to the current latest target associated object in the Object Storage Service, that is, the obj_x object on the Object Storage Service (OSS).

[0062] In an embodiment of the present specification, when generating a first joint snapshot, the metadata of the first file block in the first joint snapshot can be pointed to the current latest target associated object in the object storage service. When performing write-time copy subsequently, the metadata of the first file block can be modified to be associated with the target associated object copy. The target associated objects pointed to by the metadata of the first file block in the joint snapshots at different times are the same or different. There is no need to copy a copy of the target associated object each time a joint snapshot is generated. This not only maintains the information of all target associated objects, but also minimizes the number of associated objects in the object storage service, and reduces storage costs.

[0063] In an optional implementation of this embodiment, the method further includes:

[0064] After the second joint snapshot is generated, if a data modification instruction for the target associated object is received, the target associated object is copied to obtain a target associated object copy, and metadata of the first file block in the second joint snapshot is pointed to the target associated object copy, wherein the second joint snapshot is a snapshot generated at a time after the first joint snapshot;

[0065] Modify the target object in the object storage service based on the data modification instruction of the target object.

[0066] In actual implementation, a corresponding joint snapshot can be automatically generated at set intervals. If the target object corresponding to the first file block in the object storage service is not modified after the joint snapshot is generated, the current joint snapshot can still be associated with the latest target object in the object storage service. If the joint snapshot is modified, it is copied first and then the target object is modified.

[0067] Specifically, if after the second joint snapshot is generated, if a data modification instruction for the target-associated object corresponding to the first file block is received, the target-associated object is first copied to obtain a copy of the target-associated object, and the metadata of the first file block in the second joint snapshot is pointed to the copy of the target-associated object. Then, based on the data modification instruction for the target-associated object, the target-associated object in the object storage service is modified.

[0068] Using the above example, Figure 3 As shown, the target object corresponding to the first file block blk_x is the obj_x object on the object storage service OSS. After the joint snapshot PBD-2 is created, the data of the target object obj_x remains unchanged. Therefore, the first file block blk_x temporarily points to obj_x in PBD-2, and no separate copy of the target object is generated.

[0069] Afterwards, a joint snapshot PBD-3 is generated. After generating the joint snapshot PBD-3, the obj_x object corresponding to the first file block blk_x needs to be modified. According to the write-time copy rule, the target associated object obj_x is first copied to obtain the target associated object copy x-3. The target associated object copy x-3 is associated with PBD-3 before the obj_x object is modified.

[0070] Afterwards, a joint snapshot PBD-4 is generated. After the joint snapshot PBD-4 is generated, the obj_x object corresponding to the first file block blk_x needs to be modified. According to the write-time copy rule, the target associated object obj_x is first copied to obtain the target associated object copy x-4. The target associated object copy x-4 is associated with PBD-4 before the obj_x object is modified.

[0071] Afterwards, a joint snapshot PBD-5 is generated. After the joint snapshot PBD-5 is generated, the data of the target associated object obj_x is not modified, so the first file block blk_x in PBD-5 temporarily points to the current latest obj_x, and no separate target associated object copy is generated.

[0072] Afterwards, a joint snapshot PBD-6 is generated. After the joint snapshot PBD-6 is generated, the data of the target associated object obj_x is not modified. Therefore, the first file block blk_x in PBD-6 temporarily points to the current latest obj_x, and no separate target associated object copy is generated.

[0073] Afterwards, a joint snapshot PBD-7 is generated. After generating the joint snapshot PBD-7, the obj_x object corresponding to the first file block blk_x needs to be modified. According to the write-time copy rule, the target associated object obj_x is first copied to obtain the target associated object copy x-7. The target associated object copy x-7 is associated with PBD-7 before the obj_x object is modified.

[0074] It should be noted that after the second joint snapshot is generated, if a data modification instruction is received for the target-associated object, two steps are required. The first step is to copy the target-associated object, and then let the metadata of the first file block in the second joint snapshot point to the copied copy to ensure that when the target-associated object is subsequently modified, the data state pointed to by the metadata of the first file block in the second joint snapshot is not affected. The second joint snapshot is associated with the copy of the target-associated object before modification, not the target-associated object after modification, ensuring that the second joint snapshot can still accurately reflect the state of the data at the time of generating the snapshot; the second step is to actually modify the target-associated object in the object storage service according to the data modification instruction.

[0075] In the embodiments of this specification, through the above-mentioned write-time copy mechanism, the target associated object can be copied before the data of the target associated object is modified, and the second joint snapshot can be associated with the copy of the target associated object copied before the modification. Therefore, the modification of the target associated object will not affect the data state at the time of generating the second joint snapshot, and the integrity of the previously generated second joint snapshot is preserved, which facilitates subsequent data tracing and recovery operations and improves the efficiency and reliability of data processing in the database.

[0076] In an optional implementation of this embodiment, after copying the target associated object to obtain a copy of the target associated object and pointing the metadata of the first file block in the second joint snapshot to the copy of the target associated object, the method further includes:

[0077] The metadata of the first file block in the first joint snapshot is associated with the target associated object copy pointed to by the metadata of the first file block in the second joint snapshot.

[0078] It should be noted that after the first joint snapshot was generated, the target object pointed to by the metadata of the first file block was not modified. The metadata of the first file block in the joint snapshot temporarily pointed to the latest target object in the object storage service. Subsequently, a second joint snapshot was generated. After the second joint snapshot was generated, the target object in the object storage service was modified based on the data modification instructions for the target object. This means that the latest target object in the object storage service has changed and is no longer the target object at the time the first joint snapshot was generated.

[0079] However, since the target associated object is copied before the target associated object in the object storage service is modified, the obtained target associated object copy is associated with the metadata of the first file block in the second joint snapshot. The target associated object copy is the target associated object at the time of generating the first joint snapshot. Therefore, in actual implementation, the metadata of the first file block in the first joint snapshot is associated with the target associated object copy pointed to by the metadata of the first file block in the second joint snapshot, that is, the metadata of the first file block in the first joint snapshot points to the target associated object copy in the second joint snapshot.

[0080] Using the above example, Figure 3 As shown, the first file block blk_x in PBD-2 can be pointed to the target associated object copy x-3 in PBD-3. That is to say, the target associated objects associated with the first file block blk_x in the joint snapshots PBD-2 and PBD-3 are both the target associated object copy x-3, so that when the data of the first file block blk_x in the PBD-2 joint snapshot needs to be read subsequently, the data of the target associated object copy x-3 can be returned from the object storage service OSS.

[0081] like Figure 3As shown, the first file block blk_x in PBD-5 can be pointed to the target associated object copy x-7 in PBD-7, and the first file block blk_x in PBD-6 can also be pointed to the target associated object copy x-7 in PBD-7. That is to say, the target associated objects associated with the first file block blk_x in the joint snapshots PBD-5, PBD-6, and PBD-7 are all target associated object copy x-7, so that when the data of the first file block blk_x in the joint snapshots PBD-5, PBD-6, and PBD-7 need to be read subsequently, the data of the target associated object copy x-7 can be returned from the object storage service OSS.

[0082] In an embodiment of the present specification, the file system can support snapshots of a cloud disk storage device and a joint object storage service, and there is no need to maintain copies of all target-associated objects associated with each joint snapshot on the cloud disk storage device. Only one copy needs to be copied when the target-associated object is modified, and the metadata of the first file block in the first joint snapshot and the second joint snapshot is associated with the copy, thereby ensuring that modifications to the target-associated object after the first joint snapshot and the second joint snapshot are generated will not affect the data status at the time of generating the first joint snapshot and the second joint snapshot, retaining the integrity of the first joint snapshot and the second joint snapshot generated previously, facilitating subsequent data tracing and recovery operations, and improving the efficiency and reliability of data processing in the database.

[0083] In an optional implementation of this embodiment, the method further includes:

[0084] Set the pointer flag information of each file block in the joint snapshot at each moment in the master version joint snapshot;

[0085] Among them, the pointing flag information is used to indicate whether the corresponding file block points to an independent target-associated object copy in the joint snapshot at each moment. The pointing flag information includes a first pointing flag and a second pointing flag. The first pointing flag is used to indicate that the corresponding joint snapshot does not point to an independent target-associated object copy, and the second flag is used to indicate that the corresponding joint snapshot points to an independent target-associated object copy.

[0086] It should be noted that since there may be many joint snapshots of cloud disk storage devices and the metadata space of file blocks is relatively limited, in actual implementation, pointing flag information can be used to indicate whether the file block points to an independent target associated object copy in the joint snapshot at each moment.

[0087] Specifically, the pointing flag information may include a first pointing flag and a second pointing flag, wherein the first pointing flag is used to indicate that the corresponding joint snapshot does not point to an independent target-associated object copy, and the second flag is used to indicate that the corresponding joint snapshot points to an independent target-associated object copy. The first pointing flag and the second pointing flag may be configured based on actual scenario requirements. For example, the pointing flag information is binary information, and the first pointing flag may be set to "0" and the second pointing flag may be set to "1". If the indicator flag corresponding to a file block in a joint snapshot is "0", it indicates that the file block in the joint snapshot does not point to an independent target-associated object copy. If the indicator flag corresponding to a file block in a joint snapshot is "1", it indicates that the file block in the joint snapshot points to an independent target-associated object copy.

[0088] In actual implementation, the pointing flag information of each file block in the joint snapshot at each moment can be recorded in the main version joint snapshot. Assuming that there are joint snapshots at n moments, the pointing flag information of each file block in the target file includes an n-bit pointing flag, and each bit of the pointing flag corresponds to a joint snapshot at a moment.

[0089] Using the above example, Figure 3 As shown, the primary version joint snapshot PBD-1 can record pointing flag information. Assuming that the joint snapshots at each moment include PBD-2 through PBD-7, the pointing flag information for file block 1 is "000000," indicating that the data for file block 1 is stored on the cloud disk storage device. The pointing flag information for file block blk_x is "011001," indicating that the metadata for file block blk_x in PBD-2, PBD-5, and PBD-6 does not point to an independent target associated object replica. However, the metadata for file block blk_x in PBD-3, PBD-4, and PBD-7 points to independent target associated object replicas (x-3, x-4, and x-7). The pointing flag information for file block 3 is "000000," indicating that the data for file block 1 is stored on the cloud disk storage device. The pointer flag information for file block blk_y is "000100," indicating that the metadata for file block blk_y in PBD-2, PBD-3, PBD-4, PBD-6, and PBD-7 does not point to an independent target-associated object copy, and the metadata for file block blk_y in PBD-5 points to an independent target-associated object copy (y-5). The pointer flag information for file block blk_z is "101000," indicating that the metadata for file block blk_z in PBD-3, PBD-5, PBD-6, and PBD-7 does not point to an independent target-associated object copy, and the metadata for file block blk_z in PBD-2 and PBD-4 points to independent target-associated object copies (x-2 and x-4).

[0090] In addition, new joint snapshots can be generated periodically and the oldest joint snapshot can be deleted. Version information for each joint snapshot can also be maintained in the master version joint snapshot. For example, the minimum version number records the smallest joint snapshot version currently stored, and the maximum version number records the largest joint snapshot version currently stored. The number of version numbers from the minimum to the maximum version number corresponds to the number of bits in the pointer flag information.

[0091] Using the above example, Figure 3 As shown in the figure, [snapshot] is the version information of each joint snapshot, where min represents the minimum version number. For example, if min is "2", it means the currently stored minimum joint snapshot version is PBD-2; max represents the maximum version number. For example, if max is "7", it means the currently stored maximum joint snapshot version is PBD-7.

[0092] It should be noted that the master version joint snapshot can maintain the pointer information of each file block in the joint snapshot at each moment, and each joint snapshot can maintain the pointer information of each file block in its own joint snapshot. Figure 3 As shown, the joint snapshot PBD-2 maintains the pointer flag "0" of the file block blk_x, the pointer flag "0" of the file block blk_y, and the pointer flag "1" of the file block blk_z.

[0093] In the embodiment of this specification, the pointing flag information is used to indicate whether the file block points to an independent target associated object copy in the joint snapshot at each moment, thereby saving metadata space of the file block.

[0094] In an optional implementation of this embodiment, after the metadata of the first file block in the second joint snapshot is pointed to the target associated object copy, the following further includes:

[0095] The pointing flag corresponding to the second joint snapshot in the pointing flag information of the first file block is updated from the first pointing flag to the second pointing flag.

[0096] In actual implementation, when generating the second joint snapshot, the target-associated object is not copied, and the metadata of the first file block in the second joint snapshot temporarily points to the latest target-associated object in the object storage service. That is, the metadata of the first file block in the second joint snapshot does not point to an independent copy of the target-associated object. That is, initially, the pointing flag corresponding to the first file block in the second joint snapshot is the "first pointing flag". After generating the second joint snapshot, when modifying the target-associated object corresponding to the first file block, the target-associated object will be copied first to obtain a copy of the target-associated object, and the metadata of the first file block in the second joint snapshot will be pointed to the copy of the target-associated object. That is, at this time, the metadata of the first file block in the second joint snapshot points to an independent copy of the target-associated object. At this time, the pointing flag corresponding to the first file block in the second joint snapshot should be updated from the first pointing flag to the second pointing flag.

[0097] Using the above example, Figure 3 As shown, the target object corresponding to the first file block blk_x is the obj_x object on the object storage service OSS. After the joint snapshot PBD-2 is created, the data of the target object obj_x remains unchanged. Therefore, the first file block blk_x temporarily points to obj_x in PBD-2. No separate copy of the target object is generated, and the pointer flag of the first file block blk_x in the joint snapshot PBD-2 is "0."

[0098] Afterward, a joint snapshot PBD-3 is generated. When PBD-3 is generated, the data of the target associated object obj_x is not modified. Therefore, the first file block blk_x in PBD-3 temporarily points to obj_x. No separate target associated object copy is generated, and the pointer flag of the first file block blk_x in PBD-3 is initially set to "0." Assume that after generating PBD-3, the obj_x object corresponding to the first file block blk_x needs to be modified. First, a copy of the target associated object obj_x is made to obtain target associated object copy x-3. This target associated object copy x-3 is then associated with PBD-3, and then the obj_x object is modified. At this point, the first file block blk_x in PBD-3 points to the separate target associated object copy x-3, and the pointer flag of the first file block blk_x in PBD-3 is updated from "0" to "1."

[0099] In an embodiment of the present specification, after the metadata of the first file block in the second joint snapshot is pointed to the target associated object copy, the pointing flag corresponding to the second joint snapshot in the pointing flag information of the first file block can be updated from the first pointing flag to the second pointing flag. The pointing flag can indicate whether each joint snapshot currently points to an independent target associated object copy, thereby ensuring the reliability of the pointing flag information.

[0100] In an optional implementation of this embodiment, the method further includes:

[0101] The pointing mark information of each file block in the joint snapshot at each moment is stored in the object storage service.

[0102] It should be noted that when restoring or restoring file data at a certain moment based on a joint snapshot at that moment, a problem will be encountered. The metadata of the first file block in the joint snapshot at that moment does not record the complete pointing mark information of the joint snapshots at each moment, but only records its own pointing mark information. If the metadata of the first file block in the joint snapshot at that moment does not point to an independent target-associated object copy, then after reading the joint snapshot at that moment, based on the pointing mark of the first file block in the joint snapshot at that moment, it can only be known that the first file block is stored on the object storage service, and it cannot be known which target-associated object is associated with the first file block, so data recovery or restoration cannot be achieved.

[0103] The pointing mark information of the target associated object associated with the first file block in the joint snapshot at each moment is only maintained in the main version joint snapshot. When restoring file data based on the joint snapshot, if the joint snapshot at a certain moment has been mounted on the machine, mounting the corresponding main version joint snapshot may consume more resources.

[0104] Therefore, in actual implementation, the pointer information of each file block in the joint snapshot at each time point can be stored in the object storage service, using the object storage service's global shared storage to resolve the version information issue of the file blocks in the joint snapshot at each time point. The pointer information of the first file block is only saved to the object storage service when the first file block in the master version joint snapshot is copied on write. In this way, when reading the joint snapshot at a certain time point, if the data of a certain file block is found in the object storage service, the metadata (i.e., the pointer information) of the file block stored in the object storage service is read on the object storage service. If the pointer information is the second pointer, the corresponding target associated object copy is directly read. If the pointer information is the first pointer, the first second pointer after the first pointer is found based on the pointer information. The independent target associated object copy in the joint snapshot pointed to by the second pointer is the data to be read. If there is no second pointer after the first pointer of the joint snapshot, it means that the latest target associated object in the object storage service is the data to be obtained, realizing data recovery and restoration functions.

[0105] Using the above example, Figure 3As shown in the figure, assuming that data needs to be restored based on the joint snapshot PBD-2, the metadata of blk_x in the joint snapshot PBD-2 does not have complete version information. We only know that the data corresponding to blk_x is on the object storage service, but we cannot know which object the file block blk_x in the joint snapshot PBD-2 is associated with. This is because when the joint snapshot PBD-2 was created, blk_x had not yet started copy-on-write, and the object associated with blk_x was still obj_x at that time. However, obj_x may have been modified or deleted later (triggering copy-on-write), and it is no longer the data when the PBD-2 snapshot was created.

[0106] When reading the joint snapshot PBD-2, it is found that the data of the file block blk_x is on the object storage service. At this time, the pointing flag information "011001" recorded in the object storage service can be read to find the first second pointing flag "1" after the first pointing flag "0" of the file block blk_x. The independent target associated object copy "x-3" pointed to in the joint snapshot PBD-3 corresponding to the second pointing flag "1" is the data that needs to be read to restore the joint snapshot PBD-2.

[0107] In addition, if Figure 3 As shown in the figure, suppose you want to restore data based on the joint snapshot PBD-6. The metadata for blk_z in the joint snapshot PBD-6 does not have complete version information. We only know that the data corresponding to blk_z is in the object storage service, but we don't know which object the file block blk_z in the joint snapshot PBD-6 is associated with. When reading the joint snapshot PBD-6, we find that the data of file block blk_z is in the object storage service. At this time, we can read the pointer flag information "101000" recorded in the object storage service and find the first pointer flag "0" of file block blk_z. There is no subsequent second pointer flag "1". At this time, we can determine that file block blk_z in the joint snapshot PBD-6 points to the latest target associated object obj_z in the object storage service. This is the data needed to restore the joint snapshot PBD-6.

[0108] In the embodiments of this specification, the pointing mark information of each file block in the joint snapshot at each moment can be stored in the object storage service, and the global shared storage of the object storage service is used to solve the problem of version information of the file blocks in the joint snapshot at each moment, avoiding the need to mount the corresponding master version joint snapshot when a joint snapshot at a certain moment has already been mounted on the machine, thereby reducing resource consumption.

[0109] In an optional implementation of this embodiment, the method further includes:

[0110] In response to a data acquisition instruction at a target time, acquiring a target joint snapshot corresponding to the target time;

[0111] Determine a first file block to be migrated to the object storage service based on the target joint snapshot, and obtain pointer information of the first file block in the object storage service;

[0112] Reading the target associated object pointed to by the metadata of the first file block in the object storage service based on the pointing flag information;

[0113] A target file at a target moment is obtained based on the second file block and the target associated object in the target joint snapshot.

[0114] In actual implementation, when it is necessary to recover or restore the target file at a certain moment, a data acquisition instruction at that moment can be initiated. This moment is the target moment. Based on the target moment, the corresponding target joint snapshot can be read. If the metadata of a file block in the target joint snapshot indicates that the corresponding data is migrated to the object storage service, the pointing mark information of the first file block at each moment of the joint snapshot can be obtained in the object storage service. Based on this pointing mark information, the corresponding target-related object can be found in the object storage service, the corresponding data can be obtained, and the target file at the target moment can be obtained by combining it with the second file block stored in the cloud disk storage device.

[0115] Using the above example, Figure 3 As shown, the target joint snapshot corresponding to the target time is joint snapshot PBD-2. The metadata of file blocks blk_x and blk_z in joint snapshot PBD-2 indicates that the corresponding data has been migrated to the object storage service. The metadata of the remaining file blocks indicates that the corresponding data is stored on the cloud disk storage device PBD. The object storage service reads the pointer flag information "011001" of file block blk_x. The first second pointer flag "1" after the first pointer flag "0" of file block blk_x is found. This second pointer flag "1" points to the independent target associated object replica "x-3" in the joint snapshot PBD-3. The data of file block blk_x at the target time (i.e., target associated object replica "x-3") is obtained. The object storage service reads the pointer flag information "101000" of file block blk_z. File block blk_z points to the independent target associated object replica "z-2". The data of file block blk_z at the target time (i.e., target associated object replica "z-3") is obtained. Restore the target file at the target time based on the file blocks stored on the cloud disk storage device PBD, as well as the target-related object replica "x-3" and the target-related object replica "z-3".

[0116] In the embodiment of this specification, when it is necessary to obtain the target file at the target time, the target joint snapshot at the target time can be read. If the metadata of a file block in the target joint snapshot indicates that the corresponding data has been migrated to the object storage service, the corresponding target associated object can be found based on the pointing flag information of the joint snapshots at each time stored in the object storage service to achieve data restoration, thereby ensuring the integrity and correctness of the target file at the target time.

[0117] In an optional implementation of this embodiment, the method further includes:

[0118] In response to a snapshot deletion instruction, deleting a specified joint snapshot;

[0119] Add a pointer marker update position field to the metadata of each file block in the master version joint snapshot record, where the pointer marker update position field is used to indicate the position of the pointer marker to be updated corresponding to the specified joint snapshot;

[0120] In response to the data modification instruction of the target associated object, the pointing flag information of the first file block is updated based on the pointing flag update position field.

[0121] It's important to note that the master version of a joint snapshot maintains the version information for each joint snapshot at each point in time, as well as the pointer information for each file block in that joint snapshot at each point in time. Deleting a specific joint snapshot (such as the oldest joint snapshot) requires not only deleting all associated objects in the Object Storage service and modifying the version information for each joint snapshot at each point in time, but also modifying the pointer information for that specific joint snapshot. This can cause significant file system instability.

[0122] In actual implementation, when deleting a specific joint snapshot, the pointing flags in the pointing flag information of the specified joint snapshot are not temporarily updated. Instead, a pointing flag update location field is added to the metadata of each file block recorded in the primary version joint snapshot. This pointing flag update location field indicates the location of the pointing flag that needs to be updated after deleting the specified joint snapshot. When data is subsequently modified (i.e., copy-on-write) for the target associated object, the pointing flags in the pointing flag information of the deleted specified joint snapshot, as well as the pointing flags that need to be updated by the copy-on-write, are synchronously updated.

[0123] Using the above example, Figure 3As shown, assuming that the joint snapshots PBD-2, PBD-3, and PBD-4 need to be deleted, the initial pointing flag information of the file block blk_x is "011001", and the positions of the pointing flags corresponding to the joint snapshots PBD-3 and PBD-4 need to be updated from "1" to "0". When deleting the joint snapshots PBD-2, PBD-3, and PBD-4, they are not updated. Instead, a pointing flag update position field "last_cow_pos" is added to the metadata of the file block blk_x to indicate that the position of the pointing flag to be updated is "the position of PBD-4". If joint snapshots PBD-8, PBD-9, and PBD-10 are subsequently generated, and PBD-10 undergoes copy-on-write, the corresponding pointing flags need to be updated. At this time, the pointing flag information of the file block blk_x initially set to "011001" before the position of "PBD-4" can be reset to "0"; and, since copy-on-write occurs for the joint snapshot PBD-10 in the newly generated joint snapshots PBD-8, PBD-9, and PBD-10, the pointing flag corresponding to the joint snapshot PBD-10 is updated to "1". Since the pointing flag information is a fixed-bit number that is recycled, the pointing flag information of the file block blk_x updated after the copy-on-write occurs can be obtained as "001001", which is used to indicate whether the file blocks blk_x in PBD-5 to PBD-10 point to independent copies of the target associated object.

[0124] It should be noted that when deleting a specified joint snapshot, the pointing flag involved in the pointing flag information of the specified joint snapshot is not modified. Only the position of the pointing flag that needs to be updated is recorded. When a write-time modification occurs subsequently, the pointing flag involved in the pointing flag information of the deleted specified joint snapshot is synchronously modified based on the recorded position, and the pointing flag that needs to be updated is copied during write, thereby reducing the number of modifications, avoiding modifying the pointing flag information every time the joint snapshot is deleted, and reducing the jitter of the file system.

[0125] An embodiment of the present specification provides a data processing method, which is applied to a file system in a database, and enables the file system in the database to migrate the first file block in a target file from a cloud disk storage device to an object storage service to reduce the storage cost of the database, and can generate a joint snapshot of the cloud disk storage device. The metadata of the first file block in the joint snapshot can directly point to the corresponding target associated object in the object storage service, and a low-cost consistent joint snapshot is implemented at the file system level. Based on the joint snapshot, functions such as full backup and restore of the database can be implemented. While reducing the storage cost of the database, important functions such as full backup and restore of the database can be retained, thereby ensuring database performance.

[0126] Figure 4A schematic diagram of a joint snapshot in a database data processing method provided by an embodiment of this specification is shown. Figure 4 As shown in the figure, a target file consists of five file blocks. The dotted box on the left represents the primary version joint snapshot PBD-1. The primary version joint snapshot contains the target file's metadata, with each box representing a file block of the target file. A file block with a solid border indicates that the latest data for this file block is stored on the shared storage virtual block device. A file block with a dotted border indicates that the latest data for this file block has been migrated to the Object Storage Service (OSS) and points to the corresponding target object on OSS.

[0127] PBD-2, PBD-3, ..., PBD-7, etc. in the upper right corner are joint snapshots generated by the file system at different times.

[0128] like Figure 4 As shown, the data for file blocks X1, Z1, and W1 are stored on the PBD, while the data for file blocks Y1 and P1 are migrated to the OSS. For file block Y1, the latest target associated object on the OSS is "Object Y1." A joint snapshot, PBD-2, is generated. After the generation of joint snapshot PBD-2, "Object Y1" on the OSS is not modified. The metadata for file block Y1 in PBD-2 temporarily points to the latest target associated object, "Object Y1," on the OSS. Subsequently, a joint snapshot, PBD-3, is generated. After the generation of joint snapshot PBD-3, "Object Y1" on the OSS needs to be modified. According to the copy-on-write rules, "Object Y1" is first copied to obtain a target associated object replica, "Object Y1 Copy." This "Object Y1 Copy" is associated with PBD-3, and the metadata for file block Y1 in PBD-2 is updated to point to the associated "Object Y1 Copy" on PBD-3. Afterwards, modify "Object Y1" on OSS to obtain "Object Y2". At this time, the latest target associated object of file block Y1 in OSS is "Object Y2".

[0129] After that, a joint snapshot PBD-4 is generated. After generating the joint snapshot PBD-4, "Object Y2" in OSS needs to be modified. According to the copy-on-write rule, "Object Y2" is first copied to obtain the target associated object copy "Object Y2 copy". This "Object Y2 copy" is associated with PBD-4. Then, "Object Y2" on OSS is modified to obtain "Object Y3". At this time, the latest target associated object of file block Y1 in OSS is "Object Y3".

[0130] Afterwards, a joint snapshot PBD-5 is generated. After the generation of joint snapshot PBD-5, "Object Y3" on the OSS is not modified. The metadata of file block Y1 in PBD-5 temporarily points to the current latest target-associated object "Object Y3" on the OSS. Afterwards, a joint snapshot PBD-6 is generated. After the generation of joint snapshot PBD-6, "Object Y3" on the OSS is not modified. The metadata of file block Y1 in PBD-6 also temporarily points to the current latest target-associated object "Object Y3" on the OSS. Afterwards, a joint snapshot PBD-7 is generated. After the generation of joint snapshot PBD-7, "Object Y3" on the OSS needs to be modified. According to the write-time copy rule, "Object Y3" is first copied to obtain a target-associated object copy "Object Y3 copy", and the "Object Y3 copy" is associated with PBD-7. Afterwards, "Object Y3" on the OSS is modified to obtain "Object Y4". At this time, the current latest target-associated object of file block Y1 on the OSS is "Object Y4".

[0131] File block W1 is stored on the PBD until PBD-3 is generated (including when PBD-2 is generated). After PBD-3 is generated, file block W1 is migrated to the OSS, and the latest target associated object on the OSS is object W1. Subsequently, a joint snapshot, PBD-4, is generated. Object W1 on the OSS is not modified after PBD-4 is generated. The metadata for file block W1 in PBD-4 temporarily points to the latest target associated object, object W1, on the OSS. Subsequently, a joint snapshot, PBD-5, is generated. After PBD-5 is generated, object W1 on the OSS needs to be modified. According to the copy-on-write rules, object W1 is first copied to obtain a target associated object replica, object W1-copy. This object W1-copy is associated with PBD-5, and the metadata for file block W1 in PBD-4 is updated to point to the associated object W1-copy in PBD-5. Afterwards, "Object W1" on OSS is modified to obtain "Object W2". At this time, the latest target associated object of file block W1 in OSS is "Object W2".

[0132] Afterwards, the latest target associated object in OSS, "object W2", is migrated back to PBD. The subsequently generated PBD-6 and PBD-7 do not point to the object in OSS. That is, the data is stored on PBD, and the latest target associated object of file block W1 in OSS is empty.

[0133] For file block P1, the latest target associated object in OSS is "Object P1". A joint snapshot PBD-2 is generated. After generating joint snapshot PBD-2, "Object P1" in OSS needs to be modified. According to the copy-on-write rule, "Object P1" is first copied to obtain a target associated object copy "Object P1 copy". This "Object P1 copy" is associated with PBD-2. Then, "Object P1" on OSS is modified to obtain "Object P2". At this time, the latest target associated object of file block P1 in OSS is "Object P2".

[0134] Afterwards, a joint snapshot PBD-3 is generated. Object P2 on the OSS is not modified after the generation of joint snapshot PBD-3. The metadata of file block P1 in PBD-3 temporarily points to the latest target associated object Object P2 on the OSS.

[0135] After that, a joint snapshot PBD-4 is generated. After generating the joint snapshot PBD-4, the "object P2" in the OSS needs to be modified. According to the write-time copy rule, the "object P2" is first copied to obtain the target associated object copy "object P2 copy". The "object P2 copy" is associated with PBD-4. Then, the "object P2" on the OSS is modified to obtain "object P3". At this time, the current latest target associated object of the file block P1 in the OSS is "object P3".

[0136] Afterwards, a joint snapshot PBD-5 is generated. After the generation of joint snapshot PBD-5, "Object P3" on the OSS is not modified. The metadata of file block P1 in PBD-5 temporarily points to the current latest target-related object "Object P3" on the OSS. Afterwards, a joint snapshot PBD-6 is generated. After the generation of joint snapshot PBD-6, "Object P3" on the OSS is not modified. The metadata of file block P1 in PBD-6 also temporarily points to the current latest target-related object "Object P3" on the OSS. Afterwards, a joint snapshot PBD-7 is generated. After the generation of joint snapshot PBD-7, "Object P3" on the OSS is not modified. The metadata of file block P1 in PBD-7 also temporarily points to the current latest target-related object "Object P3" on the OSS.

[0137] The primary version joint snapshot PBD-1 can also record pointing flag information (i.e., bitmap), indicating whether the corresponding file block in the joint snapshot at each moment points to an independent target associated object copy, as well as the version information [snapshot] of each joint snapshot. Figure 4As shown in the figure, the min value in [snapshot] is 2 and the max value is 7. The joint snapshots at each time point include PBDs 2 to 7. The pointer flag information for file block X1 is "000000", indicating that the data of file block X1 is stored on the PBD. The pointer flag information for file block Y1 is "011001", indicating that the metadata of file block Y1 in PBDs 2, 5, and 6 does not point to independent target-associated object replicas. The metadata of file block Y1 in PBDs 3, 4, and 7 points to independent target-associated object replicas (Y1, Y2, and Y3). The pointer flag information for file block Z1 is "000000", indicating that the data of file block Z1 is stored on the PBD. The pointing flag information for file block W1 is "000100," indicating that the metadata for file block W1 in PBD-2, PBD-3, PBD-4, PBD-6, and PBD-7 does not point to an independent target-associated object copy, while the metadata for file block W1 in PBD-5 points to an independent target-associated object copy (W1). The pointing flag information for file block P1 is "101000," indicating that the metadata for file block P1 in PBD-3, PBD-5, PBD-6, and PBD-7 does not point to an independent target-associated object copy, while the metadata for file block P1 in PBD-2 and PBD-4 points to independent target-associated object copies (P1 and P2).

[0138] An embodiment of the present specification provides a data processing method, which is applied to a file system in a database, and enables the file system in the database to migrate the first file block in a target file from a cloud disk storage device to an object storage service to reduce the storage cost of the database, and can generate a joint snapshot of the cloud disk storage device. The metadata of the first file block in the joint snapshot can directly point to the corresponding target associated object in the object storage service, and a low-cost consistent joint snapshot is implemented at the file system level. Based on the joint snapshot, functions such as full backup and restore of the database can be implemented. While reducing the storage cost of the database, important functions such as full backup and restore of the database can be retained, thereby ensuring database performance.

[0139] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 5 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 5 As shown, the file system applied to the database includes:

[0140] a migration module 502 configured to migrate a first file block in a target file from the cloud disk storage device to the object storage service in response to a migration instruction for the first file block, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block;

[0141] The generation module 504 is configured to generate a joint snapshot of the cloud disk storage device, wherein the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to the corresponding target associated object in the object storage service, and the target associated object is the first file block stored in the object storage service.

[0142] Optionally, the generating module 504 is further configured to:

[0143] Generate a first joint snapshot, and point the metadata of the first file block in the first joint snapshot to the target associated object in the object storage service. The first joint snapshot is a snapshot generated at any time, and the metadata of the first file block in the joint snapshots at different times may point to the same or different target associated objects.

[0144] Optionally, the device further includes a copy module configured to:

[0145] After the second joint snapshot is generated, if a data modification instruction for the target associated object is received, the target associated object is copied to obtain a target associated object copy, and metadata of the first file block in the second joint snapshot is pointed to the target associated object copy, wherein the second joint snapshot is a snapshot generated at a time after the first joint snapshot;

[0146] Modify the target object in the object storage service based on the data modification instruction of the target object.

[0147] Optionally, the replication module is further configured to:

[0148] The metadata of the first file block in the first joint snapshot is associated with the target associated object copy pointed to by the metadata of the first file block in the second joint snapshot.

[0149] Optionally, the device further includes a setting module configured to:

[0150] Set the pointer flag information of each file block in the joint snapshot at each moment in the master version joint snapshot;

[0151] Among them, the pointing flag information is used to indicate whether the corresponding file block points to an independent target-associated object copy in the joint snapshot at each moment. The pointing flag information includes a first pointing flag and a second pointing flag. The first pointing flag is used to indicate that the corresponding joint snapshot does not point to an independent target-associated object copy, and the second flag is used to indicate that the corresponding joint snapshot points to an independent target-associated object copy.

[0152] Optionally, the device further includes an updating module configured to:

[0153] The pointing flag corresponding to the second joint snapshot in the pointing flag information of the first file block is updated from the first pointing flag to the second pointing flag.

[0154] Optionally, the device further includes a storage module configured to:

[0155] The pointing mark information of each file block in the joint snapshot at each moment is stored in the object storage service.

[0156] Optionally, the device further includes an acquisition module configured to:

[0157] In response to a data acquisition instruction at a target time, acquiring a target joint snapshot corresponding to the target time;

[0158] Determine a first file block to be migrated to the object storage service based on the target joint snapshot, and obtain pointer information of the first file block in the object storage service;

[0159] Reading the target associated object pointed to by the metadata of the first file block in the object storage service based on the pointing flag information;

[0160] A target file at a target moment is obtained based on the second file block and the target associated object in the target joint snapshot.

[0161] Optionally, the device further includes a deletion module configured to:

[0162] In response to a snapshot deletion instruction, deleting a specified joint snapshot;

[0163] Add a pointer marker update position field to the metadata of each file block in the master version joint snapshot record, where the pointer marker update position field is used to indicate the position of the pointer marker to be updated corresponding to the specified joint snapshot;

[0164] In response to the data modification instruction of the target associated object, the pointing flag information of the first file block is updated based on the pointing flag update position field.

[0165] One embodiment of the present specification provides a data processing device, which is applied to a file system in a database, and enables the file system in the database to migrate the first file block in a target file from a cloud disk storage device to an object storage service to reduce the storage cost of the database, and can generate a joint snapshot of the cloud disk storage device. The metadata of the first file block in the joint snapshot can directly point to the corresponding target associated object in the object storage service, and a low-cost consistent joint snapshot is implemented at the file system level. Based on the joint snapshot, functions such as full backup and restore of the database can be implemented. While reducing the storage cost of the database, important functions such as full backup and restore of the database can be retained, thereby ensuring database performance.

[0166] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.

[0167] Figure 6 6 shows a block diagram of a computing device according to one embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0168] Computing device 600 also includes an access device 640 that enables computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0169] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0170] Computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 can also be a mobile or stationary server.

[0171] The processor 620 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.

[0172] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned data processing method are of the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0173] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0174] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0175] An embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above-mentioned data processing method when executed by a processor.

[0176] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned data processing method.

[0177] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0178] Computer instructions include computer program code, which may be in source code, object code, executable files, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of computer-readable media may be appropriately expanded or reduced based on the requirements of patent practice. For example, in some jurisdictions, according to patent practice, computer-readable media does not include electric carrier signals or telecommunications signals.

[0179] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0180] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0181] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, applied to a file system in a database, comprising: In response to a migration instruction for a first file block in a target file, migrate the first file block from a cloud disk storage device to an object storage service, wherein the target file is divided into at least one file block, and the first file block is any one of the at least one file block; Generate a joint snapshot of the cloud disk storage device, wherein the joint snapshot refers to a consistent snapshot obtained by combining the data snapshot of the cloud disk storage device and the target associated object pointed to by the first file block migrated to the object storage service, the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to the corresponding target associated object in the object storage service, the target associated object is the first file block stored in the object storage service, before modifying the target associated object, copy the target associated object, associate a second joint snapshot to the copy of the target associated object copied before the modification, and then modify the target associated object, the second joint snapshot is a snapshot generated at a time after the first joint snapshot, and the first joint snapshot is a snapshot generated at any time; Among them, the method also includes: setting the pointing flag information of each file block in the joint snapshot at each moment in the master version joint snapshot, the pointing flag information is used to indicate whether the corresponding file block points to an independent target associated object copy in the joint snapshot at each moment, and the master version joint snapshot is used to persistently store the information of each joint snapshot.

2. The data processing method according to claim 1, wherein generating a joint snapshot of the cloud disk storage device comprises: Generate a first joint snapshot, and point the metadata of the first file block in the first joint snapshot to a target associated object in the object storage service, wherein the target associated objects pointed to by the metadata of the first file block in the joint snapshots at different times are the same or different.

3. The data processing method according to claim 2, further comprising: After generating the second joint snapshot, upon receiving a data modification instruction for the target associated object, copying the target associated object to obtain a copy of the target associated object, and pointing the metadata of the first file block in the second joint snapshot to the copy of the target associated object; Modify the target-related object in the object storage service based on the data modification instruction of the target-related object.

4. The data processing method according to claim 3, further comprising: after copying the target associated object to obtain the target associated object copy and pointing the metadata of the first file block in the second joint snapshot to the target associated object copy; The metadata of the first file block in the first joint snapshot is associated with the target associated object copy pointed to by the metadata of the first file block in the second joint snapshot.

5. The data processing method according to claim 3, wherein the pointing flag information includes a first pointing flag and a second pointing flag, wherein the first pointing flag is used to indicate that the corresponding joint snapshot does not point to an independent target associated object copy, and the second pointing flag is used to indicate that the corresponding joint snapshot points to an independent target associated object copy.

6. The data processing method according to claim 5, further comprising: after pointing the metadata of the first file block in the second joint snapshot to the target associated object copy; The pointing flag corresponding to the second joint snapshot in the pointing flag information of the first file block is updated from the first pointing flag to the second pointing flag.

7. The data processing method according to claim 5, further comprising: The pointing mark information of each file block in the joint snapshot at each moment is stored in the object storage service.

8. The data processing method according to claim 7, further comprising: In response to a data acquisition instruction at a target time, acquiring a target joint snapshot corresponding to the target time; Determining a first file block to be migrated to the object storage service based on the target joint snapshot, and obtaining pointing flag information of the first file block in the object storage service; Reading the target associated object pointed to by the metadata of the first file block in the object storage service based on the pointing flag information; A target file at the target moment is obtained based on the second file block in the target joint snapshot and the target associated object.

9. The data processing method according to any one of claims 1 to 8, further comprising: In response to a snapshot deletion instruction, deleting a specified joint snapshot; Adding a pointer marker update position field to the metadata of each file block recorded in the master version joint snapshot, wherein the pointer marker update position field is used to indicate the position of the pointer marker to be updated corresponding to the specified joint snapshot; In response to the data modification instruction of the target associated object, the pointing flag information of the first file block is updated based on the pointing flag update position field.

10. A data processing device, applied to a file system in a database, comprising: a migration module configured to migrate a first file block in a target file from a cloud disk storage device to an object storage service in response to a migration instruction for the first file block, wherein the target file is divided into at least one file block and the first file block is any one of the at least one file block; A generation module is configured to generate a joint snapshot of the cloud disk storage device, wherein the joint snapshot refers to a consistent snapshot obtained by combining a data snapshot of the cloud disk storage device and a target associated object pointed to by a first file block migrated to the object storage service, the joint snapshot includes metadata of each file block in the target file, the metadata of the first file block points to a corresponding target associated object in the object storage service, the target associated object is the first file block stored in the object storage service, before modifying the target associated object, the target associated object is copied, a second joint snapshot is associated with the copy of the target associated object copied before modification, and then the target associated object is modified, the second joint snapshot is a snapshot generated at a time after the first joint snapshot, and the first joint snapshot is a snapshot generated at any time; Among them, the device also includes a setting module, which is configured to: set the pointing flag information of each file block in the joint snapshot at each moment in the main version joint snapshot, and the pointing flag information is used to indicate whether the corresponding file block points to an independent target associated object copy in the joint snapshot at each moment, and the main version joint snapshot is used to persistently store the information of each joint snapshot.

11. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program / instruction, which implements the steps of the data processing method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Data processing method and device and electronic equipment

    CN114077574A

  • Snapshot processing method and system

    CN116225782A