Data management method based on distributed storage system

By responding to file access requests in the Ceph system, obtaining uncachedated shards and saving them in local cached files, the problem of low data reading efficiency caused by object device memory fragmentation is solved, and more efficient data reading and higher information security is achieved.

CN119987646APending Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311501360.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the Ceph system, the memory of the object device cannot identify whether the data of different RADOS objects belongs to the same file, resulting in data fragmentation, thereby reducing data reading efficiency.

Method used

By responding to the file access request of the target object, obtain the file description information of the target file and determine the cache result of the target file in the local cache set. If the cache result is partially cached, the uncache shard is obtained from the object storage system, and saved it in the local cache file, and returned to the target object after update.

Benefits of technology

It reduces time-consuming data positioning, improves data reading efficiency, reduces network bandwidth and transmission resources consumption, and improves data information security and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987646A_ABST
    Figure CN119987646A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data management method and device based on a distributed storage system, equipment and a storage medium, which can be applied to various scenes such as artificial intelligence, cloud technology, intelligent traffic and auxiliary driving, the method responds to a file access request triggered by a target object, obtains file description information of a target file, and sends the file description information to the storage medium. And obtaining a target cache result of the target file in the local cache set. And when the target cache result is partial cache, obtaining the non-cached fragment of the target file from the object storage system, and storing the non-cached fragment in the first cache file, so that the updated first cache file is returned to the target object, and the data reading efficiency is improved while the file access requirement of the target object is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data management method, apparatus, computing device and storage medium based on a distributed storage system. Background Art

[0002] With the continuous development of science and technology, distributed storage systems are gradually applied to many fields to improve the redundancy and availability of data storage. Among them, Ceph Distributed Storage System (Ceph) system can build a stable and efficient data storage infrastructure due to its high availability, high scalability and open source code. It is often used in many business scenarios such as cloud storage, virtualization and container storage backend, backup and archiving. It is the preferred storage solution for large-scale applications and cloud computing environments.

[0003] In the related technology, in the Ceph system, when an object expects to obtain specific data, it can use the Ceph-fuse client installed on the object device to remotely read the RADOS object data pre-stored in the object storage (Reliable AutonomicDistributed Object Store, RADOS) system, and cache the RADOS object data in the memory of the object device. In this way, when the object expects to read the same data again, it can directly obtain the data from the memory of the object device to meet the object's data access request.

[0004] However, the memory on the object device caches data at the granularity of RADOS object data, which leads to a high degree of memory fragmentation and reduces data reading efficiency.

[0005] For example, in a Ceph system, a file is usually divided into multiple file shards, each of which corresponds to one or more RADOS object data. However, the object device cannot identify whether different RADOS object data belong to the same file, so the RADOS object data of different files are mixed and stored in the memory of the object device, resulting in a high degree of data fragmentation in its memory.

[0006] Therefore, when the object device needs to read specific data from the memory, it must search all RADOS object data cached in the entire memory to determine the required target RADOS object data, resulting in a long data location time and low data reading efficiency.

[0007] In summary, how to improve data reading efficiency is an urgent problem to be solved. Summary of the invention

[0008] The present application provides a data management method, device, equipment and computer storage medium based on a distributed storage system to improve data reading efficiency.

[0009] In a first aspect, an embodiment of the present application provides a data management method based on a distributed storage system, the method comprising:

[0010] In response to a file access request triggered by a target object, obtain file description information of the target file, and based on the file description information, obtain a target cache result of the target file in a local cache set;

[0011] When the target cache result is a partial cache, based on the file description information, obtaining at least one uncached fragment corresponding to the target file from the object storage system; the partial cache indicates that the complete target file is not stored in the local cache set, and the at least one uncached fragment indicates partial data of the target file that is not stored in the local cache set;

[0012] The at least one uncached segment is saved in a first cache file, and the updated first cache file is returned to the target object; the first cache file is a local cache file in the local cache set corresponding to the target file.

[0013] In a second aspect, an embodiment of the present application provides a data management device based on a distributed storage system, the device comprising:

[0014] An acquisition unit, configured to obtain file description information of a target file in response to a file access request triggered by a target object, and obtain a target cache result of the target file in a local cache set based on the file description information;

[0015] A processing unit is configured to obtain, when the target cache result is a partial cache, at least one uncached fragment corresponding to the target file from the object storage system based on the file description information; the partial cache indicates that the complete target file is not stored in the local cache set, and the at least one uncached fragment indicates partial data of the target file that is not stored in the local cache set;

[0016] The sending unit is used to save the at least one uncached fragment in a first cache file and return the updated first cache file to the target object; the first cache file is a local cache file corresponding to the target file in the local cache set.

[0017] Optionally, the file description information includes: a target file identifier and a target storage location of the target file, and the device further includes a matching unit, configured to:

[0018] Based on the target file identifier, perform a file query on the local cache set;

[0019] When there is no target cache file corresponding to the target file identifier in the local cache set, determining that the target cache result is completely uncached;

[0020] When there is a target cache file corresponding to the file identifier in the local cache set, obtaining the file storage location of at least one cached slice corresponding to the target cache file; and performing region matching on the target storage location based on the file storage location of each cached slice;

[0021] When it is determined that there is a partial overlapping area between the target storage location and at least one file storage location, determining that the target cache result is a partial cache;

[0022] When it is determined that the target storage location completely overlaps with at least one file storage location, the target cache result is determined to be completely cached.

[0023] Optionally, the target storage location includes: a file start offset and a file end offset of the target file in the local cache set, and the file storage location corresponding to each cached fragment includes: a fragment start offset and a fragment end offset of the cached fragment in the cache file, then the matching unit is specifically used to:

[0024] For at least one file storage location, perform the following operations:

[0025] For a file storage location, compare the slice start offset corresponding to the file storage location with the file start offset and the file end offset respectively to obtain a first comparison result;

[0026] Compare the end offset of the fragment corresponding to the file storage position with the start offset of the file and the end offset of the file to obtain a second comparison result;

[0027] Based on the first comparison result and the second comparison result, a corresponding area matching result is obtained. Optionally, the matching unit is specifically configured to:

[0028] When the file start offset of the target file is less than the slice start offset, and the file end offset is greater than the slice start offset, it is determined that there is a partial overlap between the target storage location and the file storage location; or,

[0029] When the file end offset is greater than the segment end offset and the file start offset is less than the segment end offset, it is determined that there is a partial overlap area between the target storage location and the file storage location.

[0030] Optionally, the device further comprises a fusion unit, configured to:

[0031] In the local cache set, determining a first cache file corresponding to the target file, and obtaining a file storage location of each of at least one cached fragment corresponding to the first cache file;

[0032] For the at least one uncached shard, perform the following operations respectively:

[0033] For an uncached slice, when there is a matching cache slice corresponding to the uncached slice in the at least one cached slice, the matching cache slice is merged with the uncached slice and saved in the first cache file; wherein the matching cache slice is adjacent to the file storage location of the uncached slice.

[0034] Optionally, the matching cache shard is determined by a shard matching operation, and the fusion unit is specifically configured to:

[0035] For the at least one cached shard, perform the following operations respectively:

[0036] For a cached segment, performing endpoint matching on the file storage location of the cached segment and the file storage location of the uncached segment to obtain an endpoint matching result;

[0037] When it is determined based on the endpoint matching result that the file storage locations of the cached segment and the uncached segment are adjacent, the cached segment is used as the matching cache segment.

[0038] Optionally, the fusion unit is specifically used for:

[0039] If the slice start offset of the cached slice is the same as the slice end offset of the uncached slice, it is determined that the file storage locations of the cached slice and the uncached slice are adjacent;

[0040] If the end offset of the cached segment is the same as the start offset of the uncached segment, it is determined that the file storage locations of the cached segment and the uncached segment are adjacent.

[0041] Optionally, the fusion unit is specifically used for:

[0042] Based on the file start offset of the cached file in the local cache set, select the slice start offset with the smallest relative distance to the file start offset of the cached file from the slice start offset of the matching cached slice and the slice start offset of the uncached slice as the target start offset;

[0043] Based on the file end offset of the cached file, from the fragment end offset of the matching cached fragment and the fragment end offset of the uncached fragment, select the fragment end offset with the smallest relative distance to the file end offset of the cached file as the target end offset;

[0044] A target fused slice is obtained based on the target start offset and the target end offset, and the cache file is updated based on the target fused slice.

[0045] Optionally, the processing unit is further used for:

[0046] When the target cache result is that all are cached, obtaining a second cache file corresponding to the target file from the local cache set, and returning the second cache file to the target object;

[0047] When the target cache result is that all files are not cached, a complete target file is obtained from the object storage system, and the target file is returned to the target object.

[0048] Optionally, the device further includes an updating unit, configured to:

[0049] In response to a cache update instruction, obtaining a cache set identifier currently used by the local cache set;

[0050] replacing the cache set identifier of the local cache set with a preset disable identifier, and stopping using the local cache set;

[0051] Corresponding to the cache set identifier, a new local cache set is generated, and local caching is performed through the new local cache set.

[0052] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any one of the data management methods in the first aspect.

[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program is run on a computer device, the computer program is used to enable the computer device to execute any one of the data management methods in the first aspect above.

[0054] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes any one of the data management methods in the above-mentioned first aspect.

[0055] The beneficial effects of this application are as follows:

[0056] In the embodiment of the present application, the local cache data management is performed locally with the semantics of the file system layer through the local cache set, that is, the cache files of the local cache set are divided according to different cache files and the structure of the cached fragments of each cache file, so as to realize the efficient storage and management of the local cache set and reduce the degree of fragmentation of the local cache set. Based on the local cache set, the embodiment of the present application can quickly and accurately determine the cache status of the target file in the local cache set through the file description information of the target file carried by the file access request, reduce the time consumption of data positioning, and improve the efficiency of data reading. Moreover, when it is determined that the cache result of the target file is a partial cache, that is, the complete target file is not saved in the local cache set, the embodiment of the present application only needs to remotely obtain the partial data of the target file that is not saved in the local cache set from the object storage system, that is, the uncached fragments of the target file, thereby avoiding re-acquiring all the data of the target file from the object storage system, while satisfying the data access request of the target object, further reducing the unnecessary consumption of network bandwidth and transmission resources, and improving the efficiency of data reading.

[0057] Furthermore, the embodiments of the present application can store data remotely read from the object storage system in the local disk of the object device through a local cache set, thereby realizing persistent data storage, reducing the risk of data loss and damage that may be caused by restarting the object device and the Ceph-fuse client installed thereon, improving the information security and reliability of the data, and avoiding the process of having to re-read and cache related data from the object storage system after each restart of the object device or the Ceph-fuse client, thereby saving network bandwidth and transmission resources.

[0058] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0060] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0061] Figure 2 A schematic diagram of another application scenario provided by an embodiment of the present application;

[0062] Figure 3 A schematic diagram of a data reading and caching process of multiple Ceph-fuse clients in the prior art;

[0063] Figure 4 A schematic diagram of a data reading and caching process of multiple Ceph-fuse clients provided in an embodiment of the present application;

[0064] Figure 5 A system architecture diagram of a data management device provided in an embodiment of the present application;

[0065] Figure 6 A method flow chart of a data management method based on a distributed storage system provided in an embodiment of the present application;

[0066] Figure 7 A schematic diagram of the structure of a cache file in a local cache set provided in an embodiment of the present application;

[0067] Figure 8 A schematic diagram of the structure of each cache slice in a cache file provided in an embodiment of the present application;

[0068] Fig. 9 A schematic diagram of the structure of a local cache set provided in an embodiment of the present application;

[0069] Fig.10 A schematic diagram of partially overlapping regions provided in an embodiment of the present application;

[0070] FIG11( a ) is a schematic diagram of a region completely overlapping provided in an embodiment of the present application;

[0071] FIG11( b ) is a schematic diagram of a non-overlapping region provided in an embodiment of the present application;

[0072] Fig.12 A schematic diagram of a target file reading process provided in an embodiment of the present application;

[0073] Fig.13 A schematic diagram of an endpoint matching process provided in an embodiment of the present application;

[0074] Fig.14 A schematic diagram of a slice fusion process provided in an embodiment of the present application;

[0075] Fig.15 Shown is a schematic diagram of updating a local cache set provided by an embodiment of the present application;

[0076] Fig.16 A schematic diagram of a data management device based on a distributed storage system provided in an embodiment of the present application;

[0077] Fig.17 A schematic diagram of the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0078] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.

[0079] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application can mean at least two, for example, can be two, three or more, and the embodiments of the present application are not limited.

[0080] The term "and / or" in the embodiments of the present application is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0081] It is understandable that in the following specific implementations of the present application, data related to retention rate and the like are involved. When the various embodiments of the present application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, when it is necessary to obtain file data related to objects in an object storage system, relevant volunteers can be recruited and relevant agreements on volunteer authorization data can be signed, and then the data of these volunteers can be used for implementation; or, by implementing within the scope of an authorized organization, the following implementation methods can be implemented using the data of members within the organization to manage data; or, the relevant data used in the specific implementation are all simulated data, such as simulated data generated in a virtual scene.

[0082] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:

[0083] Distributed Storage System (DSS): A storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions. Distributed storage systems can distribute data across multiple physical nodes or servers to provide storage solutions with high availability, scalability, and fault tolerance.

[0084] Ceph distributed storage system: also known as Ceph system, an open source distributed storage system that can provide three storage methods: object, file, and block storage services. It has the advantages of high reliability, high scalability, and high performance. The Ceph system includes a variety of storage service systems and components, such as object storage systems and distributed file systems.

[0085] CephFS system: Stores files and directories in a distributed manner on multiple physical or virtual devices in the Ceph cluster, and uses a metadata server (MDS) to manage the metadata of the file system to provide a file system with high availability, capacity expansion, and redundancy. A file system is an organizational structure for managing files and directories, which defines how to store, access, and manage file data and metadata. The CephFS system allows users to operate files in distributed storage like a local file system.

[0086] Ceph-fuse client (Ceph Filesystem in Userspace, Ceph-fuse): is the user of the CephFS system, providing access to the CephFS system in user mode through the FUSE kernel module. The Ceph client can interact with the Ceph storage cluster through the file system interface to perform data reading and writing and metadata operations.

[0087] RADOS system: One of the core components of the Ceph system, it is an object-based, automatically managed distributed storage engine. The RADOS system can divide data into multiple objects and store and manage them on multiple storage nodes to ensure high availability and data redundancy.

[0088] Block storage service (RADOS Block Device, RBD): Block storage service provided by the Ceph system. Allows users to create and manage block devices, which can be attached to computing resources such as virtual machines, physical servers or containers to store data. RBD uses the RADOS system as the underlying storage engine to provide high-performance and reliable block storage solutions.

[0089] MDS: It is a key component of the CephFS system. It is responsible for storing and managing the directory tree structure in the CephFS system and metadata information related to files and directories, including file and directory attributes, permissions, and index information. The CephFS system can also provide file services for Ceph-fuse clients to find and access files through MDS.

[0090] Metadata: Descriptive information about files and directories, including file size, owner, permissions, creation time, modification time, etc.

[0091] Index Node (Inode): A data structure used to store file attributes and data block locations. In the Ceph system, each file data has a unique corresponding Inode, which represents the file's metadata information, such as file type, permissions, owner, creation time, size, etc. It also contains a pointer to the area where the actual data of the file is located. Therefore, in the Ceph system, Inode can be used as a metadata identifier for a file or directory to represent and manage the attributes and status of a file or directory.

[0092] Offset: represents the specific location of data in a file, usually expressed in bytes, and can be used to indicate where to start reading or writing data in the file, so as to locate and access data in the file. By setting a reasonable offset, you can achieve random access to different parts of the file without having to read the entire file in sequence.

[0093] Temporal locality: Over a period of time, data access operations tend to be concentrated on a small portion of data, that is, this portion of data will be accessed multiple times in a short period of time.

[0094] Cache: A technology for temporarily storing data, which transfers data stored on slow physical devices to high-speed memory or flash devices through some mechanism to improve the speed of data access and system throughput to accelerate access to data. In the embodiment of the present application, file data is cached through a local cache set to reduce the number of times data is read from a remote storage system (such as a Ceph storage cluster) and improve access performance.

[0095] Persistent storage: Store data or information on non-volatile storage media (such as hard disks and solid-state drives) for a long time to ensure that the data is still available after the system is shut down or the power is off. In this solution, the local cache set is used for persistent storage of file data to ensure data persistence and reliability.

[0096] Cloud technology: refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0097] The embodiments of the present application mainly relate to the cloud storage aspect of cloud technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.

[0098] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID entity, ID). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.

[0099] The process of the storage system allocating physical storage space to a logical volume is as follows: based on the estimated capacity of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of the Redundant Array of Independent Disks (RAID), the physical storage space is divided into stripes in advance. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.

[0100] The following is a brief introduction to the design concept of the embodiments of the present application.

[0101] With the continuous development of science and technology, distributed storage systems are gradually applied to many fields to improve the redundancy and availability of data storage. Among them, the Ceph system, with its high availability, high scalability and open source, can better build a stable and efficient data storage infrastructure. It is often used in many business scenarios such as cloud storage, virtualization and container storage backend, backup and archiving. It is the preferred storage solution for large-scale applications and cloud computing environments.

[0102] For example, in Ceph systems, data caching is usually performed at the granularity of RADOS object data. A file is usually divided into multiple file fragments, each of which corresponds to one or more RADOS object data, and each RADOS object has a unique object identifier. An object device with a Ceph-fuse client installed can remotely read the RADOS object data in its associated RADOS system and store it in the device memory. When the object expects to obtain the same RADOS object data again, the object device can directly read the data from the memory through the corresponding object identifier and send it to the object to satisfy the object's data access request.

[0103] However, when caching data at the granularity of RADOS object data, there are problems such as the object device being unable to identify whether different RADOS object data belong to the same file, causing RADOS object data of different files to be mixedly stored in the memory of the object device, resulting in a high degree of memory fragmentation of the object device, thereby reducing data reading efficiency.

[0104] For example, when an object needs to access specific data, the corresponding object device must search all RADOS object data cached in the entire memory to determine the required target RADOS object data, resulting in a long data location time for the object device and low data reading efficiency. In addition, once the complete data of the target RADOS object requested by the object is not stored in the memory of the object device, the entire target RADOS object must be retrieved from the object storage system, resulting in unnecessary consumption of network bandwidth and transmission resources, reducing data reading efficiency.

[0105] In order to solve the above problems, an embodiment of the present application provides a data management method based on a distributed storage system. The technical effects that can be achieved by this method will be described in detail in subsequent embodiments and will not be elaborated here.

[0106] After introducing the design ideas of the embodiments of the present application, the following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.

[0107] The solution provided in the embodiment of the present application can be applied to scenarios where Ceph distributed storage systems and distributed storage systems with Ceph-like architectures perform data storage and management. Figure 1 As shown, it is a schematic diagram of an application scenario provided by an embodiment of the present application. The application scenario diagram includes a data management device 100 and an object storage system 110.

[0108] The data management device 100 may be a computer device with certain processing capabilities, such as a tablet computer (PAD), a laptop computer, a personal computer (PC), or a server, etc., which can be configured to execute any of the methods and devices provided in the embodiments of the present application, and no further examples are given here. The data management device 100 may be installed with a client related to the Ceph system, which may be software, a web page, a small program, etc., which is not specifically limited in the embodiments of the present application.

[0109] Among them, the client can be a Ceph-fuse client, and the object can use the Ceph-fuse client installed on the data management device 100 to communicate with the associated object storage system 110. The Ceph-fuse client is mainly used to receive operation instructions input by the user, such as reading data, writing data, etc. For the convenience of description, the following takes the execution subject of the method as a server that can execute the method as an example to introduce the implementation method of the method. It can be understood that the execution subject of the method is a server is only an exemplary description and should not be understood as a limitation on the method. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, but is not limited to this. The server may include one or more processors, memories, and I / O interfaces for interacting with terminals. In addition, the server may also be configured with a database, which may be used to store each cache file in the local cache set, each cached shard, and the file storage location of the cached shard, etc. The memory of the server may also store program instructions of the data management method based on the distributed storage system provided in the embodiment of the present application, which, when executed by the processor, can be used to implement the steps of the data management method based on the distributed storage system provided in the embodiment of the present application, so as to improve data reading efficiency.

[0110] The object storage system 110 is a distributed storage system that can store and manage data, and provides data storage and access services for the data management device 100, such as the underlying storage system RADOS provided by the CephFS system. The RADOS system can divide data into one or more RADOS objects and store them in a distributed manner on multiple storage nodes or servers and other physical devices, thereby providing reliable data storage and access services.

[0111] The data management device 100 and the object storage system 110 may be directly or indirectly connected to each other through one or more networks 120. The network 120 may be a wired network or a wireless network, for example, a mobile cellular network or a wireless fidelity (Wireless-Fidelity, WIFI) network, or other possible networks, which are not limited in the embodiment of the present invention.

[0112] In a possible implementation, the data management method based on the distributed storage system in the embodiment of the present application can be executed by a data management device 100, which can be a terminal device 101 or a server 102, that is, the method can be executed by the terminal device 101 or the server 102 alone, or can be executed jointly by the terminal device 101 and the server 102.

[0113] Specifically, refer to Figure 2 The figure shows another application scenario schematic diagram provided by an embodiment of the present application, in which the terminal device 101 and the server 102 jointly execute the data management method provided by the embodiment of the present application. When the object expects to obtain specific file data by operating the Ceph-fuse client on the terminal device 101, the terminal device initiates a file access request to the server 102. The server 102 obtains the file description information of the target file based on the file access request, and obtains the target cache result of the target file in the local cache set of the server according to the file description information. If the complete target file is not saved in the local cache set, the uncached fragment corresponding to the target file is remotely obtained from the object storage system 110 according to the file description information, and saved in the local cache file corresponding to the target file, thereby returning the updated complete local cache file to the target object.

[0114] It is worth noting that in the scenario where a server with multiple Ceph-fuse clients is interacting with the object storage system, refer to Figure 3 The figure shows a schematic diagram of the data reading and caching process of multiple Ceph-fuse clients in a server under the prior art. In the prior art, the server will divide the corresponding storage space, i.e., the client memory, from its own server memory for each Ceph-fuse client to store the RADOS object data required by each Ceph-fuse client. However, under the prior art, the data in the memory of each client is independent of each other and cannot perceive each other. When multiple Ceph-fuse clients all expect to read the same RADOS object data in the RADOS system, the server will perform multiple repeated data reading and caching operations, resulting in unnecessary consumption of transmission resources. In addition, the server caches the data in the memory of each client divided from its memory. Once the server or Ceph-fuse client is restarted, it will lead to the risk of data loss and damage, reducing the information security and reliability of the data.

[0115] In view of the above problems, reference Figure 4The diagram shows a data reading and caching process of multiple Ceph-fuse clients on a server provided by an embodiment of the present application. In a scenario where multiple Ceph-fuse clients are running on a server, the server 102 executes the data management method provided by an embodiment of the present application alone, and stores the data read from RADOS by each Ceph-fuse client on the server in a local cache set on the server's local disk to achieve persistent storage of data. Even if the server or Ceph-fuse client is restarted, the server does not need to read and cache the relevant data from RADOS again, saving network bandwidth and transmission resources, avoiding the risk of data loss and damage caused by restarting the server and its installed Ceph-fuse client, and improving the information security and reliability of the data. In addition, each Ceph-fuse client on the server can share the relevant data of the local cache set stored in the server's local disk. When multiple clients all expect to access the same RADOS object data on the RADOS system, the cached cache files can be directly obtained from the local cache set, without the server having to repeatedly perform data reading and caching operations.

[0116] It should be noted that, in the embodiment of the present application, the number of data management devices 101 may be one or more, and similarly, the number of object storage systems 102 may be one or more. Figure 1-2 , 4 are only examples. In fact, the number and communication methods of data management devices and object storage systems are not limited and are not specifically limited in the embodiments of the present application.

[0117] Of course, the method provided in the embodiment of the present invention is not limited to the above Figure 1-2 , 4, can also be used in other possible application scenarios, and the embodiment of the present invention does not limit this. Figure 1-2 and Figure 4 The functions that can be implemented by each device in the application scenario shown are described in the subsequent method embodiments and will not be elaborated here.

[0118] like Figure 5 As shown, it is a system architecture diagram of a data management device provided in an embodiment of the present application. In this architecture, the following parts may be included:

[0119] (1) A file reading module is used to respond to and parse file access requests triggered by objects, send the file description information carried by the file access request to the cache management module, and read uncached data remotely from the RADOS system according to the read instructions issued by the cache management module, or directly read cached files locally and return them to the client related to the object.

[0120] (2) A file writing module is used to cache the uncached data obtained by the file reading module from the RADOS system into a local cache set, and to upload the relevant data to the RADOS system in response to a file writing request triggered by an object.

[0121] (3) A cache management module is used to manage and maintain the local cache set, implement operations such as updating cache files, cache shards, and local cache sets, and implement local persistent storage and file query functions by managing the file structure of the local cache set. The cache management module can determine the cache result corresponding to the target file in the local cache set based on the cache status of the local cache set and the file description information sent by the file reading module, and then send corresponding instructions to the file reading module based on different cache results.

[0122] (4) Local cache module, responsible for storing cache files and local storage space of cache slices.

[0123] refer to Figure 6 As shown, Figure 6 This is a method flow chart of a possible data management method based on a distributed storage system provided in an embodiment of the present application. The method can be applied to a terminal device or a server. The specific implementation process of the method is as follows:

[0124] Step 601: In response to a file access request triggered by a target object, obtain file description information of a target file.

[0125] In an embodiment of the present application, when a target object expects to access a target file stored in a CephFS system, a file access request may be triggered through a Ceph-fuse client. The file access request carries file description information of the target file, which is used to indicate data information such as the storage location and file size of the target file.

[0126] In a possible implementation, the target object may be an object in different forms such as a user or an application, and the target object may trigger a file access request according to various forms such as a user request, a scheduled task or an automated script, that is, the file access request may be initiated by the target object through a command line tool, a graphical interface operation, an application or an application programming interface (API). The target file may be a text file, an image, a video or other types of data pre-stored in the object storage system by the corresponding object, and the file access request may be a request to read the entire target file, or a request to read part of the target file, for example, to read the entire file A, or to read a file segment (data segment) with a length of 200 bytes starting from the 100th byte of file A. The file description information of the target file may include storage location information such as the index node and file path of the target file, and file size information such as the offset and file length of the target file, as well as any information that can be used to indicate the local cache status of the target file in the data management device, and the embodiment of the present application does not specifically limit this.

[0127] Specifically, take the background service process of a website as the target object as an example. The website stores website-related data through the CephFS system. The background service process of the website triggers a file access request at a fixed time, hoping to obtain the relevant file data pre-stored on the RADOS system of the website, and each file data in the RADOS system has been divided into multiple file fragments and stored in different locations. When the background service process expects to read the data at the location of file 1 fragment 5, it will generate corresponding file description information, including the Uniform Resource Locator (URL) of file 1, the offset of fragment 5 and the length of fragment 5, for example, URL = "CephFS: / / example.txt", offset = 100, len = 200. After receiving the file description information, the Ceph-fuse client will interact with the metadata server of the Ceph system, and the MDS will be responsible for parsing the URL into the index node corresponding to file 1 as Inode1, and return it to the Ceph-fuse client. The Ceph-fuse client can determine the specific data information that the target object needs to access based on the obtained index node of file 1 and the file description information such as the offset and length of shard 5 in the requested file 1, thereby executing the subsequent file reading process.

[0128] Step 602: Based on the file description information of the target file, obtain the target cache result of the target file in the local cache set.

[0129] In the embodiment of the present application, the object device stores the historically read data in the local cache set of the local disk, and caches the data in the place closest to the object to speed up the data reading efficiency. Moreover, based on the local cache set, compared with directly obtaining the entire target file from the object storage system, the embodiment of the present application can judge the cache result of the target file in the local cache set according to the file description information of the target file and the data cache situation of the local cache set, so as to decide whether it is necessary to read data remotely from the object storage system, what kind of data to read, or directly read the cached file from the local cache set and return it to the client related to the object, while satisfying the data access request of the object, and reducing the unnecessary network bandwidth and transmission resource consumption as much as possible, so as to improve the data reading efficiency. In a possible implementation, the local cache set can use a fixed area partitioning method to manage the data of the multiple cache files stored in it, so as to achieve efficient management of the cache files. The fixed area partitioning method in the embodiment of the present application refers to allocating storage space of corresponding size to each cache file in turn from the storage space of the local cache set according to the file size (size) of each cache file, so as to effectively utilize local storage resources and reduce the waste of storage resources by dividing the local storage space into areas of the size adapted to the cache file. In addition, the storage location of each cached file in the local cache set is recorded to quickly locate specific files, improving the file reading speed and response time. In summary, the fixed area division method simplifies the management and maintenance process of cached files in the local cache set, so that each cache file has a fixed storage location, reducing the degree of fragmentation and data dispersion, and reducing the complexity of data storage.

[0130] Specifically, refer to Figure 7 The figure shows a schematic diagram of the structure of a cache file in a local cache set provided by an embodiment of the present application. The local cache set records and uses the offset (off) of each cache file in the local cache set as its storage location, thereby allocating storage space of the local disk for each file read from the RADOS system for the first time through the offset and file size of the file. The offset off of each cached file in the local cache set has the following relationship:

[0131] off i =off i-1 +size i-1

[0132] Where i represents the unique corresponding file identifier of each cache file. In the embodiment of the present application, the index node number of the cache file in the RADOS system can be used to enable the cache files in the local cache set to be directly mapped to the files in the RADOS system, ensuring the consistency between the cache files and the files in the RADOS system.

[0133] i-1 represents the file identifier of the previous cache file of the cache file, that is, the i-1th cache file in the local cache set.

[0134] off i Represents the offset of the i-th cache file in the local cache set, that is, the distance from the starting storage position of the local cache set to the starting storage position of the i-th cache file, usually in bytes. i-1 Represents the offset of the i-1th cache file, that is, the offset of the previous cache file.

[0135] size represents the physical size of the cache file, that is, the storage space occupied by the cache file on the disk, usually in bytes. i-1 Represents the size of the i-1th cache file, that is, the physical size of the previous cache file.

[0136] In summary, through the above formula, the storage location of each cache file in the local cache set can be quickly located according to the offset and size of the cache file, so as to achieve efficient file reading and writing. In scenarios such as large files and random read data, it helps to improve the cache efficiency of the local cache set.

[0137] In a possible implementation, considering that the data read from the RADOS system may be a partial data segment of a file data (i.e., a file fragment), in order to further reduce the fragmentation of the local cache set, reference is made to Fig. 9 As shown, the embodiment of the present application also performs cache management on the file sharding level for each cache file in the local cache set, that is, on the basis of storing each cache file according to the fixed area division method, the file storage location of each cached file shard inside its corresponding cache file is recorded, so that according to the file storage location of each cached shard, it can be quickly determined whether the local cache set stores the complete target file, as well as the file shards that are not cached in the local cache set, thereby improving data caching and reading efficiency. Figure 8 The structure diagram of each cached slice contained in a cache file provided by an embodiment of the present application is shown. In the present application, the slice starting offset (offset) and slice length (len) of each cached slice can be used as the file storage location of the cached slice, wherein the offset and len of each cached slice are relative to the offset position inside the cache file. By recording and managing the offset and len of each cached slice, efficient cache management at the file slice level is achieved.

[0138] Specifically, refer to Fig. 9The figure shows a schematic diagram of the structure of a local cache set provided by an embodiment of the present application. The local cache set manages and maintains the file identifier Inode of each cache file, as well as the start offset offset and end offset offset of each cached fragment in each cache file, so as to divide the locally cached data according to the structure of the cache files of different index nodes and the different cached fragments contained in each cache file, and realize cache data management based on the semantics of the file system layer. And based on the data structure of the file system layer, the local cache set can also manage and maintain each cache file and cached fragment through a two-layer mapping rule. The two-layer mapping rule includes the mapping relationship between each file identifier and each cache file, and the mapping relationship between each cache fragment contained in each cache file and its own fragment start offset and fragment end offset. Through the two-layer mapping rule, the data query function can be realized.

[0139] In a possible implementation, the cache results of the target file in the local cache set may include three local cache situations: partial cache, completely uncached, and completely cached. Partial cache means that the local cache set stores partial data of the target file, but does not store the complete target file. The object device only needs to remotely read the partial data of the target file that is not stored in the local cache set from the object storage system. Complete cache means that the local cache set stores all data of the target file. The object device can directly read the local cache file corresponding to the target file from the local cache set and return it to the client related to the object. Complete uncached means that the local cache set does not store any data of the target file. The object device needs to remotely obtain all data of the target file from the object storage system.

[0140] In a possible implementation, when the file description information may include the target file identifier and the target storage location of the target file, the embodiment of the present application may perform a file query on the local cache set through the target file identifier, thereby quickly and accurately determining whether there is a cache file corresponding to the target file identifier in the local cache set, and when there is no cache file corresponding to the target file identifier in the local cache set, determining that the target cache result is completely uncached, so as to reduce the time consumption of data positioning. Moreover, when determining that there is a cache file corresponding to the target file identifier in the local cache set, the target storage location of the target file may be further matched with the file storage location of each cached segment corresponding to the cache file, so as to determine whether the target storage location of the target file and the file storage location of each cached segment have a partially overlapping area or a completely overlapping area, so as to determine whether the target cache result is partially cached or completely cached. In this way, through the relationship between the target storage location of the target file and the file storage location of the cached segment, the cache status of the target file in the local cache set can be accurately determined, so as to execute different file reading processes according to different cache results, avoid repeated data transmission and disk access operations, and improve the efficiency of data reading and caching.

[0141] In one possible implementation, in the process of performing region matching on the target file and the cached fragments in the local cache set, the cache data structure of the file system layer of the local cache set and the two-layer mapping rules can be combined to convert the region matching into a numerical comparison between offsets, thereby improving the accuracy of region matching and accurately determining the cache results of the target file that the target object needs to access in the local cache set.

[0142] Specifically, the slice start offset of each cached slice is compared with the file start and file end offsets of the target file, and the slice end offset is compared with the file start and file end offsets of the target file, to obtain various comparison results. Finally, according to the various comparison results, the relative size relationship between the above parameters is comprehensively determined to determine the specific type of the overlapping area between the target storage location and the file storage location, including partial overlap, complete overlap, no overlap, etc. Among them, when the file start offset of the target file is less than the slice start offset, and the file end offset is greater than the slice start offset, it is determined that there is a partial overlapping area between the target storage location and the file storage location. Alternatively, when the file end offset of the target file is greater than the slice end offset, and the file start offset is less than the slice end offset, it is determined that there is a partial overlapping area between the target storage location and the file storage location, that is, the cache result of the target file in the local cache set is partial cache. When the file start offset of the target file is equal to the shard start offset, and the file end offset is equal to the shard end offset, it is determined that the target storage location and the file storage location are completely overlapped, that is, the cache result of the target file in the local cache set is completely cached. When the file end offset of the target file is less than the shard start offset, or the file start offset of the target file is greater than the shard end offset, it is determined that all non-overlapping areas between the target storage location and the file storage location are completely uncached. In summary, through precise area matching processing, repeated reading of cached data is avoided, the amount of data transmission and the use of network bandwidth are reduced, thereby effectively optimizing the data access process and improving data reading efficiency.

[0143] Specifically, refer to Fig.10The diagram shows a schematic diagram of a partially overlapping area provided by an embodiment of the present application, taking the data segment of the 2500th byte to the 4500th byte of file A expected to be obtained by a file access request as an example, that is, the target file indicated by the file description information carried by the file access request is a data segment with a file start offset of 2500 and a file end offset of 4500 in the cache file A of the local cache set. According to the file storage position of the cached fragments of file A in the recorded local cache set, file A includes cached fragment 1 with a fragment start offset of 2000 and a fragment end offset of 3000, and cached fragment 2 with a fragment start offset of 4000 and a fragment end offset of 6000. For the respective file storage locations of cached fragments 1 and 2, the starting offsets of each fragment are compared with the file starting offset 2000 and the file ending offset 4000 of the target file, and the first comparison result is obtained, that is, the file starting offset of the target file is greater than the fragment starting offset of cached fragment 1, and less than the fragment starting offset of cached fragment 2; the file ending offset of the target file is greater than the fragment starting offset of cached fragment 1, and less than the fragment starting offset of cached fragment 2. At the same time, the ending offsets of each fragment are compared with the file starting offset and the file ending offset of the target file, and the second comparison result is obtained, that is, the file starting offset of the target file is less than the fragment ending offset of cached fragment 1, and less than the fragment ending offset of cached fragment 2; the file ending offset of the target file is greater than the fragment ending offset of cached fragment 1, and less than the fragment ending offset of cached fragment 2. By combining the first comparison result and the second comparison result, the corresponding area matching result can be obtained, that is, if the file start offset of the target file is less than the slice start offset of the cached slice 2, and its file end offset is greater than the slice start offset of the cached slice 2, it is determined that there is a partial overlap between the target storage position of the target file and the file storage position of the cached slice 2. If the file end offset of the target file is greater than the slice end offset of the cached slice 1, and its file start offset is less than the slice end offset of the cached slice 1, it is determined that there is a partial overlap between the target storage position of the target file and the file storage position of the cached slice 1.

[0144] Further, with reference to Figures 11(a) and 11(b), schematic diagrams are provided for the target storage location of the target file provided in the embodiments of the present application to completely overlap and completely not overlap with the file storage location of the cached slices. As shown in Figure 11(a), the file start offset of the target file is greater than or equal to the slice start offset of the cached slice 3, and its file end offset is less than or equal to the slice start offset of the cached slice 3. Then, according to the comparison result, it can be determined that the target storage location of the target file completely overlaps with the file storage location of the cached slice 3. As shown in Figure 11(b), the file start offset of the target file is greater than or equal to the slice end offset of the cached slice 4, then it can be determined that the target storage location of the target file completely does not overlap with the file storage location of the cached slice 4. The file end offset of the target file is less than or equal to the slice start offset of the cached slice 5, then it can be determined that the target storage location of the target file completely does not overlap with the file storage location of the cached slice 5.

[0145] Step 603: When the target cache result is a partial cache, at least one uncached fragment corresponding to the target file is obtained from the object storage system based on the file description information.

[0146] In the embodiment of the present application, when it is determined that only some file fragments of the target file are stored in the local cache set according to the file description information of the target file, that is, when the target cache result is a partial cache, the local cache set will remotely obtain some file fragments of the target file that have not been cached yet from the object storage system. Compared with re-obtaining the entire target file from the object storage system, the embodiment of the present application further reduces unnecessary network bandwidth and transmission resource consumption while satisfying the data access request of the object, thereby improving data reading efficiency.

[0147] Step 604: Save at least one uncached segment in the first cache file, and return the updated first cache file to the target object.

[0148] In an embodiment of the present application, after the uncached fragments of the target file are remotely obtained from the object storage system, the uncached fragments are stored in the local cache file corresponding to the target file, that is, the first cache file, so that the updated first cache file is a complete target file, and the updated first cache file is returned to the target object to meet the file access requirements of the target object.

[0149] In a possible implementation, after determining that the local cache set does not store the complete target file and remotely obtaining each uncached fragment of the target file from the object storage system, in order to further reduce the fragmentation of the cached files in the local cache set and improve the data reading and caching efficiency, the embodiment of the present application will also determine the first cache file corresponding to the target file in the local cache set, and obtain the file storage position of each cached fragment contained in the cache file inside the cache file. Thus, each uncached fragment is matched according to the file storage position of each cached fragment, and it is determined whether there is a matching cache fragment adjacent to the file storage position of the uncached fragment in each cached fragment, thereby determining the storage method of the uncached fragment in the cache text.

[0150] Specifically, when it is determined that there is no matching cached slice corresponding to each uncached slice in the cached slices of the local cache set, the uncached slice will be directly saved in the cache file according to the file storage position of the uncached slice, so as to realize the local persistent cache of the file data. On the contrary, when it is determined that there is a matching cached slice adjacent to the file storage position of the uncached slice in the cached slices of the local cache set, each uncached slice will be merged with the corresponding matching cached slice and then saved in the corresponding cache file, so as to merge multiple consecutive adjacent file slices, reduce the number of file slices in the cache file, further reduce the degree of fragmentation of the local cache set, and thus improve the efficiency of data reading and caching. In a possible implementation, the target cache result also includes all cached and all uncached. When the target cache result is all cached, that is, the local cache set stores all the data of the target file, the embodiment of the present application will obtain the second cache file corresponding to the target file from the local cache set, and the second cache file includes all the data of the target file, so the read second cache file can be directly returned to the target object to meet the data access request of the target object. When the target cache result is completely uncached, that is, the local cache set does not store any data of the target file, the embodiment of the present application will remotely read the complete target file from the object storage system and return the read target file to the target object.

[0151] Specifically, refer to Fig.12The diagram shows a target file reading process provided by an embodiment of the present application. A ceph-fuse client is installed on the object device, and the cache result of the target file in the local device can be determined through its cache management module based on the file description information carried by the file access request. When it is determined that the local cache set does not store the complete target file, the local uncached file fragments are obtained from the RADOS system. After obtaining the uncached fragments, the uncached fragments can be saved to the cache file by the file writing module, using two methods: fusing the uncached fragments with the matching cache fragments, or directly saving the uncached fragments, so that the updated cache file is returned to the target object through the file reading module. On the other hand, when the complete target file is stored in the local cache set, the cache file corresponding to the complete target file can be directly obtained from the local cache set through the file reading module and returned to the target object. When it is determined that no relevant data of any target file is stored in the local cache set, the complete target file can be directly obtained from the RADOS system to meet the file access requirements of the object.

[0152] In a possible implementation, the matching cached slices corresponding to each uncached slice can be determined by a slice matching operation, that is, for each cached slice, endpoint matching is performed with the file storage position of the uncached slice according to the respective file storage positions of the cached slices. When it is determined that the file storage positions of the cached slice and the uncached slice are adjacent based on the obtained endpoint matching results, the cached slice is used as the matching cached slice corresponding to the uncached slice. Further, the endpoint matching operation can be performed by comparing whether the slice start offset and the slice end offset of the uncached slice and the cached slice are the same to determine whether the file storage positions of the uncached slice and the cached slice are adjacent. For example, when the slice start offset of the cached slice is the same as the slice end offset of the uncached slice, or the slice end offset of the cached slice is the same as the slice start offset of the uncached slice, it is determined that the file storage positions of the cached slice and the uncached slice are adjacent. The above-mentioned endpoint matching operation of analyzing the physical location information of each file shard by comparing the offset value can quickly and accurately determine the storage location relationship between file shards without complex calculations or data structure processing, further improving data reading and caching efficiency.

[0153] Specifically, Fig.13As shown, the file storage position of the cached fragment a in the cache file A is [3, 4), and the file storage position of the cached fragment b is [5, 6), that is, the fragment start offset of the cached fragment a in the cache file A is 3k bytes, and the fragment end offset is 4k bytes, and the fragment start offset of the cached fragment b in the cache file A is 5k bytes, and the fragment end offset is 6k bytes. According to the respective file storage positions of the cached fragments a and b, endpoint matching is performed with the obtained file storage position [4, 5) of the uncached fragment c. For example, by comparing whether the fragment start offset and fragment end offset of the uncached fragment c are the same as those of the cached fragments a and b, it is determined whether the file storage positions of the uncached fragment c and the cached fragments a and b are adjacent. Specifically, since the fragment start offset of the uncached fragment c is the same as the fragment end offset of the cached fragment a, it is determined that the file storage positions of the uncached fragment c and the cached fragment a are adjacent. And, since the end offset of the uncached shard c is the same as the start offset of the cached shard b, it is determined that the file storage locations of the uncached shard c and the cached shard b are adjacent. In summary, when it is determined that the file storage locations of the uncached shard c and the cached shards a and b are adjacent based on the obtained endpoint matching results, the cached shards a and b are used as the matching cached shards corresponding to the uncached shard c.

[0154] In a possible implementation, in the process of merging a cached slice with a corresponding matching cached slice, the file start offset of the cached file in the local cache set can be used to select the slice start offset with the smallest relative distance to the file start offset of the cached file from the slice start offset of the matching cached slice and the slice start offset of the uncached slice as the target start offset. And according to the file end offset of the cached file, the slice end offset with the smallest relative distance to the file end offset of the cached file can be selected from the slice end offset of the matching cached slice and the slice end offset of the uncached slice as the target end offset. Finally, the target fused slice is obtained through the selected target start offset and target end offset, and the cached file is updated.

[0155] Specifically, refer to Fig.14As shown, after determining that the cached fragment m is the matching cached fragment corresponding to the uncached fragment c, according to the file start offset and file end offset of the cached file, the fragment start offset offset(m) with the smallest relative distance between the file start offsets is selected from the fragment start offsets offset(m) and offset(n) corresponding to the cached fragment m and the uncached fragment n, and the fragment end offsets offset(m) and offset(n), as the target start offset of the target fused fragment. And, the fragment end offset offset(n) with the smallest relative distance from the file end offset of the cached file is selected as the target end offset of the target fused fragment, so as to update the cache file through the obtained target fused fragment.

[0156] In a possible implementation, considering that each cache file in the local cache set may no longer have temporal locality due to reasons such as too long storage time, that is, it is no longer frequently accessed, and the cache files that do not have temporal locality occupy the storage space of the local cache set, which will affect data reading and caching efficiency. Therefore, in order to reduce the impact of temporal locality and improve data reading efficiency, the embodiment of the present application will also respond to the cache update instruction, replace the identifier of the local cache set currently in use with a preset disable identifier, and stop using the local cache set. And based on the identifier of the local cache set, a new local cache set is generated to perform local caching through the new local cache set. In this way, without losing the data stored in the local cache set and without affecting the object access history data, the old and no longer frequently used local cache set is actively eliminated, and the storage space that is no longer frequently used is gradually released, so as to achieve a smooth transition of the local cache set update.

[0157] Specifically, refer to Fig.15 The figure shows a schematic diagram of updating a local cache set provided by an embodiment of the present application. The ceph-fuse client on the object device has a periodic rotation module, which can issue cache update instructions at a preset period, such as updating the local cache set once every 24 hours. The cache update instruction will trigger the cache management module to rename the file name of the old local cache set currently in use with the current timestamp as the suffix name, or directly replace the identifier of the old local cache set with a disabled identifier to stop using the old local cache set for subsequent data storage, thereby actively eliminating the old and no longer frequently used local cache set without deleting the historically stored cache files and without affecting the object's access to historical data. In addition, a new local cache set is generated based on the file name of the original local cache set, so that subsequent data reading and caching operations are performed through the new local cache set, reducing the impact of time locality, improving data reading and caching efficiency, and releasing storage space that is no longer frequently used, achieving a smooth transition of local cache set updates.

[0158] See also Fig.16 Based on the same inventive concept, the embodiment of the present application further provides a data management device 160 based on a distributed storage system, the device comprising:

[0159] The acquisition unit 1601 is used to obtain file description information of the target file in response to a file access request triggered by the target object, and obtain a target cache result of the target file in the local cache set based on the file description information.

[0160] Processing unit 1602 is used to obtain at least one uncached fragment corresponding to the target file from the object storage system based on the file description information when the target cache result is a partial cache; the partial cache indicates that the complete target file is not stored in the local cache set, and the at least one uncached fragment indicates partial data of the target file that is not stored in the local cache set.

[0161] The sending unit 1603 is used to save at least one uncached segment in a first cache file and return the updated first cache file to the target object; the first cache file is a local cache file in the local cache set corresponding to the target file.

[0162] Optionally, the file description information includes a target file identifier and a target storage location of the target file, and the device further includes a matching unit 1604, which is used to:

[0163] Based on the target file identifier, perform file query on the local cache set;

[0164] When there is no target cache file corresponding to the target file identifier in the local cache set, determining that the target cache result is completely uncached;

[0165] When there is a target cache file corresponding to the file identifier in the local cache set, the file storage location of at least one cached slice corresponding to the target cache file is obtained; and based on the file storage location of each cached slice, the target storage location is regionally matched respectively;

[0166] When it is determined that there is a partial overlapping area between the target storage location and at least one file storage location, determining that the target cache result is a partial cache;

[0167] When it is determined that the target storage location completely overlaps with at least one file storage location, the target cache result is determined to be completely cached.

[0168] Optionally, the target storage location includes: a file start offset and a file end offset of the target file in the local cache set, and the file storage location corresponding to each cached fragment includes: a fragment start offset and a fragment end offset of the cached fragment in the cache file, then the matching unit 1604 is specifically used to:

[0169] For at least one file storage location, perform the following operations:

[0170] For a file storage location, compare the slice start offset corresponding to the file storage location with the file start offset and the file end offset to obtain a first comparison result;

[0171] Compare the end offset of the fragment corresponding to the file storage position with the start offset of the file and the end offset of the file to obtain a second comparison result;

[0172] Based on the first comparison result and the second comparison result, a corresponding region matching result is obtained.

[0173] Optionally, the matching unit 1604 is specifically configured to:

[0174] When the file start offset of the target file is less than the fragment start offset and the file end offset is greater than the fragment start offset, it is determined that there is a partial overlap area between the target storage location and the file storage location; or,

[0175] When the file end offset is greater than the fragment end offset and the file start offset is less than the fragment end offset, it is determined that there is a partial overlap area between the target storage location and the file storage location.

[0176] Optionally, the device further includes a fusion unit 1605, which is used to:

[0177] In the local cache set, determine a first cache file corresponding to the target file, and obtain a file storage location of each of at least one cached fragment corresponding to the first cache file;

[0178] For at least one uncached shard, perform the following operations:

[0179] For an uncached slice, when there is a matching cache slice corresponding to the uncached slice in at least one cached slice, the matching cache slice and the uncached slice are merged and saved in a first cache file; wherein the file storage locations of the matching cache slice and the uncached slice are adjacent.

[0180] Optionally, the matching cache shard is determined by a shard matching operation, and the fusion unit 1605 is specifically configured to:

[0181] For at least one cached shard, perform the following operations:

[0182] For a cached shard, perform endpoint matching on the file storage location of the cached shard and the file storage location of the uncached shard to obtain an endpoint matching result;

[0183] When it is determined based on the endpoint matching result that the file storage locations of the cached shard and the uncached shard are adjacent, the cached shard is used as the matching cached shard.

[0184] Optionally, the fusion unit 1605 is specifically configured to:

[0185] If the shard start offset of the cached shard is the same as the shard end offset of the uncached shard, it is determined that the file storage locations of the cached shard and the uncached shard are adjacent;

[0186] If the end offset of the cached shard is the same as the start offset of the uncached shard, it is determined that the file storage locations of the cached shard and the uncached shard are adjacent.

[0187] Optionally, the fusion unit 1605 is specifically configured to:

[0188] Based on the file start offset of the cached file in the local cache set, the slice start offset with the smallest relative distance to the file start offset of the cached file is selected from the slice start offsets of the matching cached slices and the slice start offsets of the uncached slices as the target start offset;

[0189] Based on the file end offset of the cached file, from the fragment end offset of the matching cached fragment and the fragment end offset of the uncached fragment, select the fragment end offset with the smallest relative distance to the file end offset of the cached file as the target end offset;

[0190] Based on the target start offset and the target end offset, the target fused shard is obtained, and based on the target fused shard, the cache file is updated.

[0191] Optionally, the processing unit 1602 is further configured to:

[0192] When the target cache result is that all the target files have been cached, a second cache file corresponding to the target file is obtained from the local cache set, and the second cache file is returned to the target object;

[0193] When the target cache result is completely uncached, the complete target file is obtained from the object storage system and the target file is returned to the target object.

[0194] Optionally, the apparatus further includes an updating unit 1606, configured to:

[0195] In response to the cache update instruction, obtaining a cache set identifier currently used by the local cache set;

[0196] Replace the cache set identifier of the local cache set with a preset disabled identifier, and stop using the local cache set;

[0197] Corresponding to the cache set identifier, a new local cache set is generated, and local caching is performed through the new local cache set.

[0198] The device can be used to execute the data management method based on the distributed storage system provided in each embodiment of the present application. Therefore, for the functions that can be implemented by each functional module of the device, please refer to the description of the aforementioned embodiments and no further details will be given.

[0199] It is worth mentioning that in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0200] See also Fig.17 Based on the same technical concept, the present application also provides a computer device. In one embodiment, the computer device can be Figure 1 The server or terminal device shown in FIG. Fig.17 As shown, it includes a memory 1701 , a communication module 1703 and one or more processors 1702 .

[0201] The memory 1701 is used to store computer programs executed by the processor 1702. The memory 1701 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0202] The memory 1701 may be a volatile memory, such as a random-access memory (RAM); the memory 1701 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1701 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1701 may be a combination of the above memories.

[0203] The processor 1702 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1702 is used to implement the above-mentioned data management method based on the distributed storage system when calling the computer program stored in the memory 1701 .

[0204] The communication module 1703 is used to communicate with terminal devices or other servers.

[0205] The specific connection medium between the memory 1701, the communication module 1703 and the processor 1702 is not limited in the embodiment of the present application. Fig.17 In the embodiment, the memory 1701 and the processor 1702 are connected via a bus 1704. The bus 1704 is connected to the processor 1702 via a bus 1704. Fig.17 The connections between the other components are described with bold lines, which are only for illustration and are not intended to be limiting. The bus 1704 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig.17 The diagram shows that only one thick line is used, but this does not mean that there is only one bus or only one type of bus.

[0206] The memory 1701 stores computer storage media, which stores computer executable instructions. The computer executable instructions are used to implement the data management method based on a distributed storage system in the embodiments of the present application. The processor 1702 is used to execute the data management method based on a distributed storage system in the above-mentioned embodiments.

[0207] Based on the same inventive concept, an embodiment of the present application also provides a storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes the steps of the data management method based on a distributed storage system according to various exemplary embodiments of the present application described above in this specification.

[0208] In some possible implementations, various aspects of the data management method based on a distributed storage system provided by the present application may also be implemented in the form of a computer program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to enable the computer device to execute the steps of the data management method based on a distributed storage system according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the steps of each embodiment.

[0209] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0210] The program product of the embodiment of the present application may adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can be run on a computer device. However, the program product of the present application is not limited thereto. In the present application, the readable storage medium may be any tangible medium containing or storing a program, and the computer program included therein may be used by or in combination with a command execution system, apparatus or device.

[0211] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.

[0212] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0213] Computer programs for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" language or similar programming languages.

[0214] It should be noted that, although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into multiple units to be embodied.

[0215] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0216] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0217] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0218] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A data management method based on a distributed storage system, characterized in that: The method comprises: In response to a file access request triggered by a target object, obtain file description information of the target file, and based on the file description information, obtain a target cache result of the target file in a local cache set; When the target cache result is a partial cache, based on the file description information, obtaining at least one uncached fragment corresponding to the target file from the object storage system; the partial cache indicates that the complete target file is not stored in the local cache set, and the at least one uncached fragment indicates partial data of the target file that is not stored in the local cache set; The at least one uncached segment is saved in a first cache file, and the updated first cache file is returned to the target object; the first cache file is a local cache file in the local cache set corresponding to the target file.

2. The method according to claim 1, characterized in that The file description information includes: a target file identifier and a target storage location of the target file; Then, obtaining the target cache result of the target file in the local cache set based on the file description information includes: Based on the target file identifier, perform a file query on the local cache set; When there is no target cache file corresponding to the target file identifier in the local cache set, determining that the target cache result is completely uncached; When there is a target cache file corresponding to the file identifier in the local cache set, obtaining the file storage location of at least one cached slice corresponding to the target cache file; and performing region matching on the target storage location based on the file storage location of each cached slice; When it is determined that there is a partial overlapping area between the target storage location and at least one file storage location, determining that the target cache result is a partial cache; When it is determined that the target storage location completely overlaps with at least one file storage location, the target cache result is determined to be completely cached.

3. The method according to claim 2, characterized in that The target storage location includes: the file start offset and the file end offset of the target file in the local cache set, and the file storage location corresponding to each cached fragment includes: the fragment start offset and the fragment end offset of the cached fragment in the cache file; Then, based on the respective file storage locations of the cached slices, the target storage locations are respectively region matched, including: For at least one file storage location, perform the following operations: For a file storage location, compare the slice start offset corresponding to the file storage location with the file start offset and the file end offset respectively to obtain a first comparison result; Compare the end offset of the fragment corresponding to the file storage position with the start offset of the file and the end offset of the file to obtain a second comparison result; Based on the first comparison result and the second comparison result, a corresponding area matching result is obtained.

4. The method according to claim 3, characterized in that Obtaining a corresponding region matching result based on the first comparison result and the second comparison result includes: When the file start offset of the target file is less than the slice start offset, and the file end offset is greater than the slice start offset, it is determined that there is a partial overlap between the target storage location and the file storage location; or, When the file end offset is greater than the segment end offset and the file start offset is less than the segment end offset, it is determined that there is a partial overlap area between the target storage location and the file storage location.

5. The method according to claim 1, characterized in that The step of storing the at least one uncached segment in the first cache file comprises: In the local cache set, determining a first cache file corresponding to the target file, and obtaining a file storage location of each of at least one cached fragment corresponding to the first cache file; For the at least one uncached shard, perform the following operations respectively: For an uncached slice, when there is a matching cache slice corresponding to the uncached slice in the at least one cached slice, the matching cache slice is merged with the uncached slice and saved in the first cache file; wherein the matching cache slice is adjacent to the file storage location of the uncached slice.

6. The method according to claim 5, characterized in that The matching cache shard is determined by a shard matching operation, and the shard matching operation includes: For the at least one cached shard, perform the following operations respectively: For a cached segment, performing endpoint matching on the file storage location of the cached segment and the file storage location of the uncached segment to obtain an endpoint matching result; When it is determined based on the endpoint matching result that the file storage locations of the cached segment and the uncached segment are adjacent, the cached segment is used as the matching cache segment.

7. The method according to claim 6, characterized in that The determining that the cached segment is adjacent to the file storage location of the uncached segment includes: If the slice start offset of the cached slice is the same as the slice end offset of the uncached slice, it is determined that the file storage locations of the cached slice and the uncached slice are adjacent; If the end offset of the cached segment is the same as the start offset of the uncached segment, it is determined that the file storage locations of the cached segment and the uncached segment are adjacent.

8. The method according to claim 5, characterized in that The merging of the matching cached fragment and the uncached fragment and storing the fragments in the cache file includes: Based on the file start offset of the cached file in the local cache set, select the slice start offset with the smallest relative distance to the file start offset of the cached file from the slice start offset of the matching cached slice and the slice start offset of the uncached slice as the target start offset; Based on the file end offset of the cached file, from the fragment end offset of the matching cached fragment and the fragment end offset of the uncached fragment, select the fragment end offset with the smallest relative distance to the file end offset of the cached file as the target end offset; A target fused slice is obtained based on the target start offset and the target end offset, and the cache file is updated based on the target fused slice.

9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: When the target cache result is that all are cached, obtaining a second cache file corresponding to the target file from the local cache set, and returning the second cache file to the target object; When the target cache result is that all files are not cached, a complete target file is obtained from the object storage system, and the target file is returned to the target object.

10. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: In response to a cache update instruction, obtaining a cache set identifier currently used by the local cache set; replacing the cache set identifier of the local cache set with a preset disable identifier, and stopping using the local cache set; Corresponding to the cache set identifier, a new local cache set is generated, and local caching is performed through the new local cache set.

11. A data management device based on a distributed storage system, characterized in that: The data management device comprises: An acquisition unit, in response to a file access request triggered by a target object, obtains file description information of a target file, and based on the file description information, obtains a target cache result of the target file in a local cache set; A processing unit is configured to obtain, when the target cache result is a partial cache, at least one uncached fragment corresponding to the target file from the object storage system based on the file description information; the partial cache indicates that the complete target file is not stored in the local cache set, and the at least one uncached fragment indicates partial data of the target file that is not stored in the local cache set; The sending unit is used to save the at least one uncached fragment in a first cache file and return the updated first cache file to the target object; the first cache file is a local cache file corresponding to the target file in the local cache set.

12. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to call a computer program stored in the memory, and execute the method according to any one of claims 1 to 8 according to the obtained program.

13. A computer-readable non-volatile storage medium, characterized in that: The computer-readable non-volatile storage medium stores a program, and when the program is run on a computer, the computer is enabled to implement the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The method comprises a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device executes the method according to any one of claims 1 to 10.