Data object storage method and apparatus, computer device, and storage medium
Patent Information
- Application Number
- CN202210621695.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-06-02
AI Technical Summary
然而,在数据对象存入到存储桶后,立即读到的数据未必是最新数据,可能会导致数据读取存在不完整或不准确的问题,这样无法满足写后读的应用场景
[0034]上述数据对象存储方法、装置、计算机设备、存储介质和计算机程序产品,在临时目录下创建元数据级的临时文件;将目标数据的数据块并发向存储容器上传,获得目标数据的索引信息;将目标数据的索引信息与临时文件关联,得到临时数据文件,因此若在上传过程中出现异常,数据均停留在临时目录下。此外,依据目标数据对应的文件路径在对象目录下创建第一文件目录;将临时数据文件移动至第一文件目录下,得到第一文件目录下的目标数据文件,因此即便在上传过程中出现异常,数据均停留在临时目录下,从而确保了在对象目录下的文件数据是完整的,而且满足数据强一致性的要求,在进行数据读取时,可以有效保证数据的完整性和准确性。
Smart Images

Figure CN117215477B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of object storage technology, and in particular to a data object storage method, apparatus, computer device, storage medium, and computer program product. Background Technology
[0002] In the context of big data, file storage can be inefficient due to slow upload and download speeds caused by large data volumes. Object storage solutions have emerged to address this issue. Object storage primarily separates data and metadata, storing them as data objects, such as within storage buckets. However, the data accessed immediately after a data object is stored in a bucket may not be the most up-to-date, potentially leading to incomplete or inaccurate data retrieval. This makes it unsuitable for read-after-write applications. Summary of the Invention
[0003] Therefore, it is necessary to provide a data object storage method, apparatus, computer device, computer-readable storage medium, and computer program product to address the above-mentioned technical problems, which can ensure that the stored data can be read immediately and guarantee strong data consistency.
[0004] Firstly, this application provides a data object storage method. The method includes:
[0005] Create temporary files at the metadata level in the temporary directory;
[0006] The target data blocks are uploaded concurrently to the storage container to obtain the index information of the target data;
[0007] The index information of the target data is associated with the temporary file to obtain a temporary data file;
[0008] Create a first file directory under the object directory based on the file path corresponding to the target data;
[0009] The temporary data file is moved to the first file directory to obtain the target data file in the first file directory.
[0010] Secondly, this application also provides a data object storage device. The device includes:
[0011] The first creation module is used to create metadata-level temporary files in the temporary directory;
[0012] The upload module is used to concurrently upload data blocks of the target data to the storage container and obtain the index information of the target data;
[0013] The association module is used to associate the index information of the target data with the temporary file to obtain a temporary data file;
[0014] The second creation module is used to create a first file directory under the object directory based on the file path corresponding to the target data;
[0015] The moving module is used to move the temporary data file to the first file directory to obtain the target data file in the first file directory.
[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0017] Create temporary files at the metadata level in the temporary directory;
[0018] The target data blocks are uploaded concurrently to the storage container to obtain the index information of the target data;
[0019] The index information of the target data is associated with the temporary file to obtain a temporary data file;
[0020] Create a first file directory under the object directory based on the file path corresponding to the target data;
[0021] The temporary data file is moved to the first file directory to obtain the target data file in the first file directory.
[0022] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0023] Create temporary files at the metadata level in the temporary directory;
[0024] The target data blocks are uploaded concurrently to the storage container to obtain the index information of the target data;
[0025] The index information of the target data is associated with the temporary file to obtain a temporary data file;
[0026] Create a first file directory under the object directory based on the file path corresponding to the target data;
[0027] The temporary data file is moved to the first file directory to obtain the target data file in the first file directory.
[0028] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0029] Create temporary files at the metadata level in the temporary directory;
[0030] The target data blocks are uploaded concurrently to the storage container to obtain the index information of the target data;
[0031] The index information of the target data is associated with the temporary file to obtain a temporary data file;
[0032] Create a first file directory under the object directory based on the file path corresponding to the target data;
[0033] The temporary data file is moved to the first file directory to obtain the target data file in the first file directory.
[0034] The aforementioned data object storage method, apparatus, computer equipment, storage medium, and computer program product create metadata-level temporary files in a temporary directory; concurrently upload data blocks of the target data to the storage container to obtain the index information of the target data; associate the index information of the target data with the temporary files to obtain temporary data files. Therefore, if an anomaly occurs during the upload process, the data remains in the temporary directory. Furthermore, a first file directory is created in the object directory based on the file path corresponding to the target data; the temporary data files are moved to the first file directory to obtain the target data files in the first file directory. Therefore, even if an anomaly occurs during the upload process, the data remains in the temporary directory, thus ensuring that the file data in the object directory is complete and meets the requirement of strong data consistency. This effectively guarantees the integrity and accuracy of the data during data reading. Attached Figure Description
[0035] Figure 1 This is an application environment diagram of a data object storage method in one embodiment;
[0036] Figure 2 This is a flowchart illustrating a data object storage method in one embodiment;
[0037] Figure 3 This is a schematic diagram of the directory tree structure in one embodiment;
[0038] Figure 4 This is a schematic diagram of the directory tree structure in another embodiment;
[0039] Figure 5 This is a schematic diagram illustrating the movement of temporary data files from a temporary directory to an object directory in one embodiment;
[0040] Figure 6 This is a schematic diagram illustrating the movement of temporary data files from the temporary directory to the object directory in another embodiment;
[0041] Figure 7 This is a schematic diagram of the process of reading a target data file based on index information in one embodiment;
[0042] Figure 8 This is a flowchart illustrating a data object storage method in another embodiment;
[0043] Figure 9 This is a schematic diagram of the chunked upload process in one embodiment;
[0044] Figure 10 This is a schematic diagram of the chunked upload process in another embodiment;
[0045] Figure 11 This is a schematic diagram of the state transition process for a target file in one embodiment;
[0046] Figure 12 This is a schematic diagram illustrating the logical relationship between MPU files, Part files, and Blocks, as well as data uploading, in one embodiment.
[0047] Figure 13 This is a structural block diagram of a data object storage device in one embodiment;
[0048] Figure 14 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0050] Before describing the solution of this application, the technical terms and concepts involved in this application will be explained as follows:
[0051] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of storage devices of various types (storage devices are also called storage nodes) in the network to work together through application software or application interfaces to provide data storage and business access functions to the outside world.
[0052] Currently, the storage method of storage systems is as follows: Logical volumes are created. During the creation of a logical volume, physical storage space is allocated to each logical volume. This physical storage space may consist of a single storage device or the disks of several storage devices. Clients store data on a logical volume, which means storing the data on the file system. The file system divides the data into many parts, each part being an object. Each object contains not only the data but also additional information such as a data identifier (ID, ID entity). The file system writes each object to the physical storage space of that logical volume and records the storage location information of each object. Therefore, when a client requests access to data, the file system can allow the client to access the data based on the storage location information of each object.
[0053] The process by which a storage system allocates physical storage space to a logical volume is as follows: the physical storage space is pre-divided into strips according to the capacity estimate of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of Redundant Array of Independent Disks (RAID). A logical volume can be understood as a strip, thus allocating physical storage space to the logical volume.
[0054] Cloud HDFS (Hadoop Distributed File System) can be a cloud distributed file system (or simply distributed file system) that evolves from HDFS. It can provide high-performance, highly consistent, and hierarchical file system metadata services, while utilizing low-cost and scalable object storage as data storage.
[0055] Target data can be various types of data obtained from the logical induction of objective things, such as text, images, videos, and audio.
[0056] Metadata can refer to data that describes data, and is used to indicate storage location, resource attributes, and file records, etc.
[0057] Data consistency refers to the consistency of data values across multiple nodes in a distributed file system.
[0058] Strong consistency means that after data is modified or written, the modified or written data can be obtained immediately.
[0059] Weak consistency means that after data is modified or written, it is not guaranteed that the modified or written data can be obtained immediately.
[0060] Eventual consistency is a specific manifestation of weak consistency, where data can eventually be obtained after it has been modified or written.
[0061] In one embodiment, the data object storage method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 sends an upload request to server 104. Server 104 creates a metadata-level temporary file in a temporary directory; concurrently uploads the data blocks of the target data to the storage container to obtain the index information of the target data; associates the index information of the target data with the temporary file to obtain a temporary data file; creates a first file directory in the object directory based on the file path corresponding to the temporary data file; and moves the temporary data file to the first file directory to obtain the target data file in the first file directory.
[0062] The terminal 102 can run a client for reading and writing data objects and a distributed file system webpage. This terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, IoT device, or portable wearable device. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices, etc.
[0063] Server 104 can be a standalone physical server or a server cluster consisting of multiple physical servers. It can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0064] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.
[0065] In one embodiment, such as Figure 2 As shown, a data object storage method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0066] S202, Creates metadata-level temporary files in the temporary directory.
[0067] Temporary directories can be temporary directories under the root directory of the file system, used to temporarily store metadata-level files. In a distributed file system, directories at all levels and their corresponding files can be organized and managed using inodes. These inodes are connected to form a directory tree, such as a temporary directory tree. It should be noted that inodes can store not only temporary files containing metadata, but also other metadata such as file size in bytes, owner, read, write, and execute permissions, and file processing timestamps. These timestamps include modification timestamps (mtime) and read timestamps (atime), where mtime indicates the last time the file content was modified, and atime indicates the last time the file was opened.
[0068] The temporary directory and the temporary files within it are stored as part of the inodes in the temporary directory tree. For example... Figure 3 As shown, the two nodes / Tmp and / Tmp / tmp_file can form a temporary directory tree. / Tmp can represent a temporary directory, while / Tmp / tmp_file represents the tmp_file file under the / Tmp directory.
[0069] Metadata level can be the level of metadata, and metadata is attribute information used to describe data.
[0070] Temporary files can refer to temporarily stored file information at the metadata level. For example, they can be identifiers used to represent files or data, such as at least one of file type, icon, and file name. In a distributed file system, these temporary files can be organized and managed as inodes, such as temporary file nodes being child nodes under a temporary directory. Figure 3 As shown.
[0071] In one embodiment, prior to S202, the server receives a storage request carrying target data from the terminal, and then calls a file creation interface to create a metadata-level temporary file in the temporary directory. For example, the server calls the CreateFile(pid, name) interface, using the parent node identifier (pid) and node name (name) from the CreateFile(pid, name) interface to create a temporary file (tmp_file) in the / Tmp directory. (See reference...) Figure 3 .
[0072] In one embodiment, prior to S202, the server creates a temporary directory and an object directory in the root directory of the file system. The temporary directory is configured to be invisible to the user object that is uploading the target data; that is, the temporary directory and the temporary files under it are not visually displayed on the operation page of the distributed file system, and therefore are invisible to the user object. On the other hand, the object directory is configured to be visible to the user object that is uploading the target data; that is, the object directory and its subdirectories and corresponding files are visually displayed on the operation page of the distributed file system, and therefore are visible to the user object.
[0073] Specifically, the server calls the directory creation interface, which creates a temporary directory and an object directory in the root directory of the file system. For example, the server calls the MkDir(pid, name) interface, which creates a Tmp directory and a User directory in the root directory of the file system. The Tmp and User directories can be found in [reference needed]. Figure 3 .
[0074] S204, concurrently upload the data blocks of the target data to the storage container to obtain the index information of the target data.
[0075] In this context, "target data" can refer to various types of data that need to be uploaded, including text, images, videos, and audio. A data block can be a small subset of the target data. A storage container can refer to a container in object storage used to store data objects, such as a bucket.
[0076] In one embodiment, the server divides the target data into at least two data blocks and uploads the two data blocks concurrently to a storage container to obtain the index information of the target data. For example, the server divides an image sent by a terminal into multiple image blocks; then, it writes the multiple image blocks concurrently to a storage bucket.
[0077] Specifically, the server divides the target data into blocks; uploads each data block concurrently to the storage container, and determines the block offset value corresponding to each data block during the upload process; and determines the index information of the target data based on the block offset value and the block size.
[0078] The block offset value can refer to the read offset value of a data block, such as the offset value of the first data block in the storage container when reading data blocks.
[0079] For example, suppose there are 5 data blocks, 1 to 5, each 1 megabyte (M). Data blocks 1 to 5 are written into the storage bucket in sequence. The index information of the target data is calculated based on the block identifier, block size, and read offset of the data block. For example, the index information of the target data = block identifier of the data block × block size + read offset of the data block (i.e., the offset of the first data block in the storage container).
[0080] In one embodiment, the server stores the block offset and block size in a data block list. In addition, it may also store index information in the data block list so that when data needs to be read, the block offset and block size, or the index information, can be obtained to read the corresponding target data file.
[0081] S206, associate the index information of the target data with the temporary file to obtain the temporary data file.
[0082] The temporary data file can be a data file containing target data and metadata, and the temporary data file can exist in the storage container in the form of a data object.
[0083] In one embodiment, within a temporary directory tree, temporary file nodes are associated with the index information of the target data, thereby establishing a connection between the target data's index information and the temporary files, resulting in a temporary data file. Since the temporary files contain metadata-level file information, associating the target data's index information with the temporary files links the metadata to the target data, resulting in a temporary data file containing both the target data and its metadata.
[0084] In one embodiment, S206 may specifically include: when all data blocks of the target data are concurrently uploaded to the storage container, the server associates the index information of the target data with temporary files. When at least one data block of the target data encounters an error during the upload process, the association of the index information of the target data with temporary files is stopped; and temporary files created in the temporary directory are cleaned up.
[0085] Because the move operations (such as the Rename operation) in this application are atomic, even if the upload fails, the file will remain in the temporary directory. Only the successfully uploaded file will reach the object directory and will be a complete file. The files that failed to upload in the temporary directory will be cleaned up periodically by the cleanup module.
[0086] S208, create the first file directory in the object directory according to the file path corresponding to the target data.
[0087] The file path can refer to the path where a user object stores data in a distributed file system. For example, if a user object wants to store a file named 'file' in the directory ' / dir0 / dir1', the file path would be ' / dir0 / dir1 / file'.
[0088] In one embodiment, after receiving a storage request from a terminal via a client or a distributed file system webpage, the server parses the file path corresponding to the target data from the storage request and creates a first file directory under the object directory based on that file path. The first file directory can refer to a directory of files (or simply a file directory), including file directories at various levels created under the object directory.
[0089] It's important to note that in a distributed file system, directories at all levels and their corresponding files are organized and managed using inodes. These inodes (including object directory nodes and file directory nodes at all levels) are connected to form a directory tree. Therefore, the inodes corresponding to an object directory and its file directories at all levels can be combined to form an object directory tree, which is visible to the user object. When the object directory tree and a temporary directory tree are combined, a large directory tree can be formed.
[0090] Specifically, the server generates a parent node identifier and node name based on the file path, and then calls the directory creation interface, using the parent node identifier and node name as parameters in the directory creation interface. The directory creation interface is then used to recursively create file directory nodes under the object directory node. That is, a first-level file directory node is created under the object directory node, then a second-level file directory node is created under the first-level file directory node, and so on, until all file directory nodes are created.
[0091] For example, when the file path ` / dir0 / dir1 / file` is obtained, the server generates the PID and node name based on this file path. Since `dir0` in the file path needs to be mounted to the object directory node, the parent node corresponding to `dir0` is the object directory node. The identifier of the object directory node is recorded as 1, i.e., PID = 1, and the node name is `dir0`. For `dir1` in the file path, it needs to be mounted to the node `dir0`, so we can obtain the PID corresponding to `dir1` as 2, and the node name is `dir1`. When obtaining the PID and node name `name`, the `MkDir(pid, name)` interface is called. Based on this `MkDir(pid, name)` interface, file directory nodes are recursively created. (See reference...) Figure 4 .
[0092] S210, move the temporary data file to the first file directory to obtain the target data file in the first file directory.
[0093] The target data file can refer to a user-visible data file located in the first file directory, containing the target data and corresponding metadata. The operation of moving the temporary data file is atomic; therefore, regardless of any exception encountered in any of the above steps, the file remains in the temporary directory. Only upon successful upload will it be moved to the first file directory of the object directory, ensuring that the file moved to the first file directory of the object directory is complete. Furthermore, all data read operations are strongly consistent reads, meaning that the target data file can be immediately accessed after being written to (i.e., uploaded), thus ensuring strong data consistency.
[0094] In one embodiment, the server checks whether the target data file exists in the first file directory. If the target data file does not exist in the first file directory, the server renames the temporary data file to obtain the target data file. The server then moves the target data file to the first file directory to obtain the target data file in the first file directory.
[0095] Specifically, the server calls the rename interface and passes parameters to the rename interface, including the parent node identifier and node name corresponding to the temporary data file, as well as the target parent node identifier and target node name; based on the parent node identifier and node name corresponding to the temporary data file, as well as the target parent node identifier and target node name, the temporary data file is renamed to obtain the target data file; the target data file is moved to the first file directory to obtain the target data file in the first file directory.
[0096] For example, the server calls the `Rename(source_pid, source_name, destination_pid, destination_name)` interface to obtain the parent node name (i.e., `source_pid`) and node name (i.e., `source_name`) of the temporary data file, as well as the target parent node name (i.e., `destination_pid`) and target node name (i.e., `destination_name`). Based on `source_pid`, `source_name`, `destination_pid`, and `destination_name`, the server renames the temporary data file, changing its name from `tmp_file` to the target data file (`file`), and moves `file` to the directory `User / dir0 / dir1`. Figure 5 As shown.
[0097] In one embodiment, the server checks if a data file with a version lower than the target version exists in the first file directory, corresponding to the temporary data file. If such a data file exists in the first file directory, the server updates the data files in the first file directory based on the temporary data file to obtain the target data file. If a data file with a version higher than the target version exists in the first file directory, the server refuses to move the target data file to the first file directory; wherein, the target version is the version corresponding to the temporary data file.
[0098] For example, such as Figure 6 As shown, if you want to move `tmp_file` with version number 99, the server checks if there is a file with a version number lower than 99 in the directory ` / User / dir0 / dir1`. Since the file in ` / User / dir0 / dir1` has a version number of 100, there is no file with a version number lower than 99. Therefore, the `tmp_file` with version number 99 is not renamed and moved to the ` / User / dir0 / dir1` directory. If you want to move `tmp_file` with version number 101, the server checks if there is a file with a version number lower than 101 in the directory ` / User / dir0 / dir1`. Since the file in ` / User / dir0 / dir1` has a version number of 100, there is a file with a version number lower than 101. Therefore, the `tmp_file` with version number 101 needs to be renamed and moved to the ` / User / dir0 / dir1` directory.
[0099] In the above embodiments, a metadata-level temporary file is created in the temporary directory; data blocks of the target data are concurrently uploaded to the storage container to obtain the index information of the target data; the index information of the target data is associated with the temporary file to obtain a temporary data file. Therefore, if an error occurs during the upload process, the data remains in the temporary directory. Furthermore, a first file directory is created in the object directory based on the file path corresponding to the target data; the temporary data file is moved to the first file directory to obtain the target data file in the first file directory. Therefore, if an error occurs during the upload process, the data remains in the temporary directory, thus ensuring the integrity of the file data in the object directory. Moreover, since the individual interface operations of cloud HDFS meet the requirement of strong data consistency, the integrity and accuracy of the data can be effectively guaranteed during data reading.
[0100] In one embodiment, after S210, the method further includes:
[0101] S702 receives a read request for the target data file.
[0102] In one embodiment, the server receives a read request for a target data file from a client, or receives the read request from a distributed file system webpage.
[0103] S704, in response to a read request for the target data file, reads the block size and block offset of the data block in the target data file.
[0104] The block offset value mentioned above can refer to the read offset value of a data block, such as the offset value of the first data block in the storage container when reading a data block.
[0105] In one embodiment, after receiving a read request from a client or a distributed file system webpage, the server reads the block size and block offset of the data block in the target data file from the data block list.
[0106] S706, determine the index information of the target data file based on the block size and block offset of the data block.
[0107] In one embodiment, when the number of data blocks is greater than or equal to two, the server can determine the block identifier, block size, and block offset of each data block in the target data file, and then determine the index information of the target data file based on the block identifier, block size, and block offset. For example, assuming there are five data blocks (1 to 5), each 1 MB in size, data blocks 1 to 5 are written into a storage bucket sequentially. The index information of the target data file is calculated based on the block identifier, block size, and block offset of each data block, such as: target data file index information = block identifier × block size + block offset.
[0108] S708 reads the target data file from the storage container based on the index information of the target data file.
[0109] In one embodiment, the server reads the target data file from the storage container of the object storage device based on the index information of the target data file.
[0110] To better understand the solutions in the above embodiments, the following is combined with... Figure 8 The explanation is as follows:
[0111] A server can include a data server and a metadata server. The metadata server is used to manage metadata, while the data server is used to read and write data. For example... Figure 8As shown, when a user object wants to read a video file for playback, the client sends a read request for the video file to a data server that has deployed a distributed file system. Upon receiving the read request, the data server generates an information retrieval request and sends it to the metadata server. The metadata server responds to the information retrieval request, obtains the block size and block offset of each data block in the video file, and then returns the block size and block offset to the data server. The data server calculates the index information of the video file based on the block identifier, block size, and block offset, and then sends this index information and the read request to the object storage device, or encapsulates the index information in the read request and sends the read request to the object storage device. The object storage device reads the corresponding video file from the storage bucket according to the index information and returns it to the data server. After receiving the video file, the data server returns it to the client, allowing the client to play the video.
[0112] In one embodiment, such as Figure 9 As shown, large amounts of data can be uploaded in chunks. The specific steps include:
[0113] S902 creates a metadata-level temporary block file in the temporary directory.
[0114] Temporary chunk files can refer to temporary multi-part upload (MPU) files. In a distributed file system, temporary directories and corresponding temporary chunk files can be organized and managed as inodes, which, when connected together, form a temporary directory tree.
[0115] In one embodiment, prior to S902, the server receives a storage request carrying media data from the terminal, and then calls a file creation interface to create a metadata-level temporary block file in a temporary directory. For example, the server calls the CreateFile(pid, name) interface, using the parent node identifier and node name in the CreateFile(pid, name) interface to create a temporary block file (tmp_file) in the / Tmp directory. (See reference...) Figure 10 .
[0116] S904, generate an upload identifier, and associate the upload identifier and the file path of the media data with the temporary block file to obtain the associated file.
[0117] The upload identifier (upload_id) can be used to uniquely identify temporary chunk files during the upload process. Media data includes large multimedia data, such as video files larger than 1GB. The file path of media data can refer to the path where the user object stores data in the distributed file system. For example, if the user object wants to store "file" in the directory " / test", the file path would be " / test / file".
[0118] For example, in addition to creating a temporary MPU file using the CreateFile(pid,name) interface, the server also generates a globally unique upload_id, and then associates the upload_id and the file path of the media data together with the temporary MPU file.
[0119] S906: When media data is uploaded concurrently to other storage containers in the form of fragment files, each fragment file is associated with an associated file to obtain a temporary block data file.
[0120] Among them, the part file can be a secondary file obtained by dividing the media data into at least two segments.
[0121] In one embodiment, the server calls at least two interfaces for uploading part files, uploads part files based on the called interfaces, and associates each fragment file with an associated file when media data is concurrently uploaded to other storage containers in the form of fragment files to obtain temporary block data files.
[0122] In addition, when media data is concurrently uploaded to other storage containers as fragment files, the server creates a fragment directory (Part directory) in the root directory of the file system to manage temporary fragmented data files.
[0123] In one embodiment, S906 may specifically include: when all segment files of media data are concurrently uploaded to the storage container, the server associates each segment file with an associated file. When at least one segment file of media data encounters an error during the upload process, the association of each segment file with the associated file is stopped, and the temporary block files created in the temporary directory are cleaned up.
[0124] Because the move operation in this application is atomic, even if the upload fails, the file will remain in the temporary directory. Only the successfully uploaded file will reach the object directory and will be a complete file. The files that failed to upload in the temporary directory will be cleaned up periodically by the cleanup module.
[0125] In one embodiment, when media data is uploaded concurrently to other storage containers as fragment files, the server determines the first file offset value of each fragment file; adds a fragment directory to the root directory of the file system to store each fragment file; and adds a fragment file table to store the first file offset values.
[0126] S908: Based on the file path corresponding to the temporary block data file, create a second file directory under the object directory.
[0127] In one embodiment, after receiving a storage request from a terminal via a client or a distributed file system webpage, the server parses the file path corresponding to the media data from the storage request and creates a second file directory under the object directory based on that file path. The second file directory may include file directories at various levels created under the object directory.
[0128] It's important to note that in a distributed file system, directories at all levels and their corresponding files are organized and managed using inodes. These inodes (including object directory nodes and file directory nodes at all levels) are connected to form a directory tree. Therefore, an object directory and the inodes corresponding to its file directories at all levels can be combined to form an object directory tree, which is visible to the user object. When the object directory tree, temporary directory tree, and fragment directory tree are combined, a large directory tree can be formed.
[0129] S910, move the temporary block data file to the second file directory to obtain the target block data file in the second file directory.
[0130] The target block data file can refer to objectified media data stored in blocks in a distributed file system.
[0131] In one embodiment, S910 may specifically include: the server moving the temporary block data file corresponding to the target fragment file to the second file directory, wherein the target fragment file belongs to at least one file among the fragment files, thereby obtaining the target block data file in the second file directory. Furthermore, the server determines the file offset value corresponding to the target block data file based on the first file offset value of the target fragment file.
[0132] In one embodiment, after S910, the target block file can be read. The specific steps include: the server receiving a read request for the target block data file; in response to the read request for the target block data file, reading a first file offset value from the fragment file table and determining a second file offset value corresponding to the fragment file; determining the index information of the target block data file based on the first file offset value and the second file offset value; and reading the target block data file in other storage containers based on the index information of the target block data file.
[0133] In one embodiment, the step of determining the second file shift value corresponding to the fragment file may specifically include: the server obtaining the block size and block offset value of each data block in the fragment file; and determining the second file shift value corresponding to the fragment file based on the block size and block offset value of each data block in the fragment file.
[0134] To gain a clearer and more intuitive understanding of the above solutions, we will combine them here. Figure 10 The explanation is as follows:
[0135] (1) Multi-segment upload initialization (InitMPU) phase.
[0136] The InitMPU phase corresponds to S902 to S904 mentioned above. The server calls the CreateFile(pid, name) interface to create a temporary MPU file, such as... Figure 10 The `tmp_file_0` is mounted to the temporary directory ` / Tmp`. The inode identifiers for the root directory, ` / User`, ` / Part`, ` / Tmp`, and ` / Tmp / tmp_file_0` are set from 1 to 5, thus creating an inode table. Additionally, a unique `upload_id` (e.g., `xxxxxx`) is generated and associated with the file path (`path = / test / file`) to the temporary MPU file. For example, a separate `mpu` table can be used to associate the temporary MPU file's `upload_id` and file path; conversely, a `part` table can be used to associate the MPU's part files.
[0137] (2) Fragment file upload (UploadPart) stage.
[0138] The UploadPart stage corresponds to S906 to S908 mentioned above. During the UploadPart stage, the server creates a subdirectory under the part directory to store the part files, such as... Figure 10As shown, create a subdirectory named xxxxxx under the / Part directory, and then store the part files in the xxxxxx subdirectory, such as file_1, file_2, and file_3. The corresponding part table and inode table are as follows. Figure 10 As shown.
[0139] (3) Complete the chunked upload (CompleteMPU) stage.
[0140] The CompleteMPU stage corresponds to S910 mentioned above. User objects can select part files to move, for example, choosing to move file_2 and file_3 to the / User / test directory, while not moving file_1 to the / User / test directory. This allows for selective uploading of some part files to the distributed file system for storage, resulting in a file of size 7M in the / User / test directory.
[0141] In the above embodiments, for uploading large amounts of data, a chunked upload method can be used, which can effectively improve data upload efficiency and reduce upload time. Furthermore, users can selectively upload segments of interest, increasing upload flexibility.
[0142] In one embodiment, the method further includes: the server calling the object augmentation interface or object truncation interface of the distributed file system; uploading augmentation data for the target data file to the storage container based on the object augmentation interface or object truncation interface, or truncation processing of the target data file in the storage container; or uploading augmentation data for the target block data file to other storage containers based on the object augmentation interface or object truncation interface, or truncation processing of the target block data file in other storage containers.
[0143] Supplementary data can be data added to the target data file or the target block data file. For example, for an editable data file, data can be added to the data file, such as adding a piece of text content to a Word document. This added data is called supplementary data.
[0144] In addition, truncation can refer to cutting a target data file or a target chunked data file into at least two segments, such as truncating an uploaded video or audio (such as music) to obtain at least two truncated files.
[0145] It should be noted that both the augmentation and truncation operations meet the requirements of strong data consistency, meaning that after augmentation or truncation, the contents of the target data file or target block data file can be read immediately.
[0146] For example, the server calls the AppendObject interface to add data to an editable file, or calls the TruncateObject interface to truncate a target data file.
[0147] In the above embodiments, the object supplementation interface allows for data supplementation of the target data file or target data file individually. Data supplementation can be performed directly within the distributed file system, eliminating the need to download and re-upload the complete target data file or target data file, thus reducing the complexity of data supplementation and improving data processing efficiency. Furthermore, the object truncation interface allows for truncation of the target data file or target data file individually. Again, data truncation can be performed directly within the distributed file system, eliminating the need to download and re-upload the complete target data file or target data file, further reducing the complexity of data supplementation and improving data processing efficiency.
[0148] In one embodiment, such as Figure 11 As shown, the method also includes:
[0149] S1102, search for target files that meet the state transition conditions based on the object directory tree.
[0150] The object directory tree is a directory tree composed of object directories and file directories. The target file can be the file for which the state migration task performs task management operations, including the target data file and other uploaded data files. The state migration task includes one of the following: archiving task, deletion task, and warm-up task. The archiving task is used to migrate the target file from its initial state to an archived state, such as from a standard state to an archived state; the deletion task is used to migrate the target file from its initial state to a deletion state, such as from an archived state to a deletion state; and the warm-up task is used to migrate data from its initial state to a warm-up state, such as from an archived state to a warm-up state.
[0151] The state transition condition can refer to the fact that the most recent modification timestamp (mtime) and read timestamp (atime) of the target file exceed the preset time threshold. For example, if the most recent modification timestamp of the target file exceeds one week or one month, then the target file meets the state transition condition.
[0152] In one embodiment, in response to a state timed transition task, the target file path is obtained, and the object directory tree is scanned based on the target file path to obtain candidate files; among the candidate files, the target file that meets the state transition conditions is determined.
[0153] S1104 marks the target file as an intermediate state that matches the state transition conditions.
[0154] In this context, intermediate states refer to the states between the initial state and the target state that the target file needs to migrate to. For example, if the state migration condition is an archive migration condition, then the target state that matches the archive migration condition is the archive state, and the intermediate state is the archive state; if the state migration condition is a delete migration condition, then the target state that matches the delete migration condition is the delete state, and the intermediate state is the delete state; if the state migration condition is a reheat migration condition, then the target state that matches the reheat migration condition is the reheat state, and the intermediate state is the reheat state.
[0155] Specifically, after the server finds a target file that meets the state transition conditions, it determines an intermediate state that matches the state transition conditions and marks the target file's state as the determined intermediate state. The process of marking the target file's state as the determined intermediate state can involve obtaining the target file's initial state and modifying the obtained initial state to the determined intermediate state.
[0156] S1106, After marking is completed, obtain the data list for each data block in the target file.
[0157] The data list can be either a list of data blocks or a list of data objects. A target file can contain at least one data block, each data block corresponds to at least one data segment, and each data segment can be a data object. In other words, a target file contains at least one data object.
[0158] Specifically, after marking the target file as an intermediate state matching the state transition conditions, the server retrieves a list of data objects for each data object in the target file. For example, when the target file is a file that can be archived, the server retrieves a list of data objects corresponding to each file that can be archived.
[0159] S1108, perform data processing on the data blocks in the data list.
[0160] Specifically, data processing involves migrating the storage locations of various data objects, including at least one of archiving, deletion, and reactivation. Archiving refers to migrating the data object from its initial storage location to a target storage location; reactivation involves creating a copy of the data object and storing it in the target storage location, with the copy's retention time matching the retention time specified in the reactivation task; deletion refers to removing the data object from its initial storage location.
[0161] S1110: After data processing is completed, the target file is migrated from the intermediate state to the target state.
[0162] The target state is the state that the data state migration target of the state timed migration task is to achieve. The target state includes at least one of the following: archived state, deleted state, and warm-up state.
[0163] Specifically, after the server completes the data processing of each data object in the data object list of the target file, it determines the target state corresponding to the completed data processing and modifies the state of the target file from the intermediate state to the target state.
[0164] In one embodiment, S1110 specifically includes: after completing at least one of file processing, deletion processing, and reheat processing, updating the intermediate state to the target state in the inode corresponding to the target file.
[0165] The aforementioned data state migration method involves the server searching for target files that meet the state migration conditions based on the object directory tree. The target file is marked as an intermediate state that matches the state migration conditions. After marking, a list of data objects for each data object in the target file is obtained, and data processing is performed on the data objects in the list. This ensures that the data objects contained in the target file are processed synchronously. After all data objects have been processed, the target file is migrated from the intermediate state to the target state. This achieves synchronous data state migration of the data objects in the target file and ensures the correctness of the target file after the data state migration.
[0166] In one embodiment, the method further includes: the server receiving a deletion request for a target file directory under the object directory; in response to the deletion request, calling the directory deletion interface of the distributed file system; and deleting the target file directory under the object directory based on the directory deletion interface.
[0167] The above deletion operation meets the requirement of strong data consistency. That is, after performing the deletion operation on the target data file or target block data file, the target data file or target block data file can be deleted immediately, and there is no need to search for the directories to be deleted one by one and then delete them one by one, which improves the efficiency of directory deletion.
[0168] As an example, we will use Cloud HDFS as an example to illustrate this. First, we will introduce the existing interface capabilities in Cloud HDFS and the S3 interface that needs to be implemented. Then, we will introduce how to complete the logic of ordinary upload based on the existing interface. Next, we will introduce the core design of MPU block upload. Finally, we will introduce some new features that are supported by Cloud HDFS as the underlying storage, which are helpful for big data scenarios and are outside the S3 standard.
[0169] (1) Cloud HDFS and S3
[0170] Cloud HDFS organizes all metadata information of the file system using inodes as entities. Write operations rely on a high-performance database, while read operations are accelerated using memory. An inode contains attributes such as inode_id, pid, name, and file_mode. The inode_id uniquely identifies the current inode, the pid represents the parent node's inode_id, the name represents the inode name, and the file_mode indicates whether the inode is a file or a directory. This allows for the construction of a complete directory tree from all inode entries. Furthermore, during metadata path lookups, it can recursively search for target inode information based on the file path (pid, name), thus implementing the upper-level file system interface. As a truly distributed file system, it supports the following interfaces:
[0171] 1. Metadata write operation
[0172] ●CreateFile(pid,name): Creates a file.
[0173] ●MkDir(pid,name): Creates a directory.
[0174] ●DeleteInode(pid,name): Deletes a file or directory.
[0175] ●Rename(source_pid,source_name,destination_pid,destination_name): Renames a file or directory.
[0176] ●CommitFileBlocks(inode_id,file_block_array): Modifies the data block information of the file (the file length remains unchanged or increases).
[0177] ●TruncateFile(inode_id, target_size): Truncates the file (reduces the file size).
[0178] 2. Metadata read operation
[0179] ●GetInode(path): Gets file or directory attributes.
[0180] ●ListChildInodes(pid): Retrieves the child nodes in the directory.
[0181] ●GetFileBlocks(inode_id): Retrieves information about the data blocks of a file.
[0182] It should be noted that the write operations described above are atomic, and the read operations are all strongly consistent reads.
[0183] Although S3 offers a rich set of interfaces, in big data scenarios, only the following core data read / write interfaces need to be implemented, as follows:
[0184] 3. Normal Upload
[0185] ●PutObject(key,object): Upload object (key can be understood as file path).
[0186] ●GetObject(key): Retrieves the object.
[0187] ●DeleteObject(key): Deletes the object.
[0188] ●CopyObject(source_key, destination_key)&DeleteObject(source_key): Simulates a Rename operation.
[0189] ●HeadObject(key): View the object's metadata.
[0190] ●GetBucket(key_prefix): This is equivalent to ListObjects, which retrieves some (specified prefix) or all objects within a bucket.
[0191] 4. Upload in chunks
[0192] ●InitMPU(key): Initializes a chunked upload event.
[0193] ●UploadPart(key, upload_id, part_number, part): Uploads data in chunks according to the specified chunk upload event.
[0194] ●CompleteMPU(key, upload_id, final_part_number_array): Completes the entire chunk upload event after all chunks have been uploaded (it can complete a part).
[0195] ●ListParts(key, upload_id): Retrieves the uploaded blocks from the specified block upload event.
[0196] ●ListMPUs: Retrieves ongoing chunk upload events within the bucket.
[0197] (2) Normal upload
[0198] Strong consistency in normal uploads requires that the file specified by the path be visible immediately after a successful PutObject, whether it is a GetObject, HeadObject, or ListObject.
[0199] This application references the rename mechanism used in big data jobs, dividing the file system root directory into a temporary Tmp directory and a user directory. The user directory is visible to the user. PutObject is completed in five steps: CreateFile -> chunked concurrent data upload to COS -> CommitFileBlocks -> recursively MkDir -> Rename. Figure 5 As shown.
[0200] Because the Rename operation is atomic, the file will remain in the Tmp directory regardless of which step fails. Only successfully uploaded files will reach the User directory. Therefore, only complete files are uploaded. The failed files in the Tmp directory are cleaned up periodically by the cleanup module.
[0201] Besides addressing data integrity issues, concurrent uploads also need to be considered. For PutObject, concurrent uploads of files at the same path will all return success. The final file order depends on the underlying definition of the object storage, which may depend on the physical or logical clock at a certain step, but data integrity must be guaranteed. A unique ID generation algorithm during temporary file creation ensures filenames in the Tmp directory do not conflict. Therefore, the challenge of concurrent uploads lies in Rename. In a normal file system, Rename returns a "target file already exists" error if the target file already exists. However, in the PutObject simulation, Rename must return success and has overwrite semantics. Therefore, this application adds a RenameOverwrite(source_inode_id, destination_pid, destination_name) implementation based on Rename, introducing the concept of a version number, similar to a logical clock, which is always incremented. It is generated when the temporary file is created and placed in the file's metadata. Then, during Rename, the version number is compared. If the condition that the target file already exists and its version number is smaller than the temporary file's is met, then it is overwritten; otherwise, it is not overwritten. Figure 6 As shown in the diagram. RenameOverwrite also exhibits atomicity and plays a crucial role in the MPU's chunked upload logic.
[0202] In the regular upload interface, besides PutObject, other interfaces can also be implemented through the cloud HDFS interface, as shown in Table 1:
[0203] Table 1
[0204]
[0205] (3) Upload in chunks
[0206] To accelerate the upload of large files, cloud HDFS also has the concept of data chunking, but the chunk length is fixed. S3, on the other hand, uses variable-length chunks, and the length of each chunk can be unequal. For example, assuming a 11MB file is being uploaded, cloud HDFS will split the 11MB file into 4MB chunk 0, 4MB chunk 1, and 3MB chunk 2; while S3 could use 5MB chunk 1, 5MB chunk 2, and 1MB chunk 3, or 5MB chunk 1 and 6MB chunk 2. This variable-length S3 chunking upload cannot be implemented using existing file system chunk sizes. Therefore, this application adopts a new design, using variable-length files as indexes, such as... Figure 12 As shown.
[0207] Next, we will compare chunked uploads and normal uploads. The differences between the two are shown in Table 2 below:
[0208] Table 2
[0209]
[0210] Furthermore, this application features a more refined design for MPU chunked upload, as shown in Table 3 below:
[0211] Table 3
[0212]
[0213]
[0214] It is important to emphasize that the additional logic in CreateFile or RenameOverwrite must be atomic along with itself, which is guaranteed by database transactions at the underlying level.
[0215] Next, combine Figure 10 The following describes how to calculate the length of an MPU file and how to read MPU file data:
[0216] This application uses a separate MPU table to associate the upload_id and file path of the MPU file in the InitMPU stage, and a Part table to associate the Part files of the MPU in the UploadPart stage, and completes the final calculation of the file length in the CompleteMPU stage. Figure 10 As shown.
[0217] from Figure 10 As can be seen, in the InitMPU stage, the MPU file length is 0 by default; in the UploadPart stage, the MPU file length remains 0, and the MPU file offset corresponding to the Part is 0 by default; in the CompleteMPU stage, the MPU file length is calculated to be 7M, and its valid Parts (parts 2 and 3 are retained) are preserved and the MPU file offset corresponding to the Part is updated.
[0218] For data reading, the file type must first be identified—Normal or MPU—and then different data indexes are obtained. If the file is of Normal type, the Block table is read directly and returned; if it is of MPU type, the Part table is read first, then the Block table is read and returned. The formula for calculating the target data index is as follows:
[0219] Normal file read offset (i.e., Normal file index information) = Block table block_id × block size + Block read offset;
[0220] MPU file read offset (i.e., MPU file index information) = Part table file_offset + Part file read offset;
[0221] Part file read Offset = Block table block_id × block size + Block read Offset;
[0222] Because the parts are independent of each other, the database concurrency advantage can be used to speed up the process of calculating the MPU file length and reading file data.
[0223] (3)New features
[0224] S3 is implemented using cloud HDFS as the underlying storage, supporting several new features that are very useful for big data scenarios, such as:
[0225] 1. Supports file append and truncate. S3 does not have the AppendObject and TruncateObject interfaces. When simulating file append and truncate, the complete file needs to be downloaded to the local machine first and then uploaded, which is very time-consuming. Since cloud HDFS supports random writes, AppendObject and TruncateObject are easy to implement.
[0226] 2. Supports file atime and mtime updates. S3 only has mtime, and it is updated when the file is first uploaded or overwritten. It is not the same as the atime and mtime concept in the file system. In contrast, cloud HDFS is a true file system, and its atime and mtime updates are accurate and have reference value.
[0227] 3. Instant directory deletion: S3 requires listing all objects prefixed with the directory path before deleting a directory, and then deleting them one by one. In contrast, cloud HDFS deletes directories in milliseconds, instantly.
[0228] (4) The following beneficial effects can be achieved through the embodiments of this application:
[0229] This application can solve the problem of eventual consistency in object storage, satisfying the application scenario of read-after-write, and also brings some other advantages, as follows:
[0230] 1. Functional advantages: Supports strong data consistency, has a complete directory hierarchy, supports file append and truncate, and supports file atime and mtime updates.
[0231] 2. Performance advantages: Supports millisecond-level atomic rename operations, instant directory deletion, and high-frequency directory listing.
[0232] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0233] Based on the same inventive concept, this application also provides a data object storage device for implementing the data object storage method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more data object storage device embodiments provided below can be found in the limitations of the data object storage method described above, and will not be repeated here.
[0234] In one embodiment, such as Figure 13 As shown, a data object storage device is provided, including: a first creation module 1302, an upload module 1304, an association module 1306, a second creation module 1308, and a movement module 1310, wherein:
[0235] The first creation module 1302 is used to create metadata-level temporary files in the temporary directory;
[0236] Upload module 1304 is used to concurrently upload data blocks of the target data to the storage container and obtain the index information of the target data;
[0237] The association module 1306 is used to associate the index information of the target data with a temporary file to obtain a temporary data file;
[0238] The second creation module 1308 is used to create a first file directory under the object directory based on the file path corresponding to the target data;
[0239] The moving module 1310 is used to move the temporary data file to the first file directory to obtain the target data file in the first file directory.
[0240] In one embodiment, the device further includes:
[0241] A new module has been added to create a temporary directory and an object directory in the root directory of the file system. The temporary directory is configured to be invisible to user objects that are uploading target data, while the object directory is configured to be visible to user objects that are uploading target data.
[0242] In one embodiment, the upload module is further configured to slice the target data into blocks; upload each data block concurrently to the storage container, and determine the block offset value corresponding to each data block during the upload process; and determine the index information of the target data based on the block offset value and the block size.
[0243] In one embodiment, the device further includes:
[0244] This association module is used to associate the index information of the target data with temporary files when all data blocks of the target data are uploaded to the storage container concurrently.
[0245] The stop module is used to stop associating the index information of the target data with the temporary file when at least one data block of the target data encounters an error during the upload process.
[0246] The cleanup module is used to clean up temporary files created in the temporary directory.
[0247] In one embodiment, the moving module is further configured to, when the target data file does not exist in the first file directory, rename the temporary data file to obtain the target data file; and move the target data file to the first file directory to obtain the target data file in the first file directory.
[0248] In one embodiment, the device further includes:
[0249] The update module is used to update the data files in the first file directory based on the temporary data file when a data file with a version lower than the target version exists in the first file directory, so as to obtain the target data file in the first file directory.
[0250] The stop module is used to refuse to move the target data file to the first file directory if a data file with a higher version than the target data file already exists in the first file directory; where the target version is the version corresponding to the temporary data file.
[0251] In one embodiment, the device further includes:
[0252] The receiving module is used to receive read requests for the target data file;
[0253] The read module is used to read the block size and block offset of data blocks in the target data file in response to a read request from the target data file.
[0254] The determination module is used to determine the index information of the target data file based on the block size and block offset of the data block; and to read the target data file from the storage container based on the index information of the target data file.
[0255] In the above embodiments, a metadata-level temporary file is created in the temporary directory; data blocks of the target data are concurrently uploaded to the storage container to obtain the index information of the target data; the index information of the target data is associated with the temporary file to obtain a temporary data file. Therefore, if an anomaly occurs during the upload process, the data remains in the temporary directory. Furthermore, a first file directory is created in the object directory based on the file path corresponding to the target data; the temporary data file is moved to the first file directory to obtain the target data file in the first file directory. Therefore, if an anomaly occurs during the upload process, the data remains in the temporary directory, thus ensuring the integrity of the file data in the object directory. Moreover, since the interface operations of cloud HDFS meet the requirements of strong data consistency, the integrity and accuracy of the data can be effectively guaranteed during data reading.
[0256] In one embodiment, the device further includes:
[0257] Create a module to create metadata-level temporary chunk files in a temporary directory;
[0258] The generation module is used to generate an upload identifier and associate the upload identifier and the file path of the media data with a temporary block file to obtain the associated file;
[0259] The association module is also used to associate each fragment file with an associated file to obtain a temporary block data file when media data is uploaded concurrently to other storage containers in the form of fragment files;
[0260] The second creation module is also used to create a second file directory under the object directory based on the file path corresponding to the temporary block data file;
[0261] The move module is also used to move the temporary block data file to the second file directory to obtain the target block data file in the second file directory.
[0262] In one embodiment, the device further includes:
[0263] The determination module is used to determine the first file offset value of each fragment file when media data is uploaded concurrently to other storage containers in the form of fragment files;
[0264] A new module is added to the root directory of the file system to store fragment files; a new fragment file table is added to store the first file offset value.
[0265] In one embodiment, the moving module is further configured to move the temporary block data file corresponding to the target fragment file to the second file directory, wherein the target fragment file belongs to at least one of the fragment files;
[0266] The determination module is used to determine the file offset value corresponding to the target block data file based on the first file offset value of the target fragment file.
[0267] In one embodiment, the device further includes:
[0268] The receiving module is used to receive read requests for the target block data file;
[0269] The read module is used to respond to a read request for the target chunk data file by reading the first file offset value from the fragment file table and determining the second file offset value corresponding to the fragment file.
[0270] The determination module is used to determine the index information of the target block data file based on the first file offset value and the second file offset value;
[0271] The read module is also used to read target block data files in other storage containers based on the index information of the target block data file.
[0272] In one embodiment, the determining module is further configured to obtain the block size and block offset value of each data block in the fragment file; and determine the second file offset value corresponding to the fragment file based on the block size and block offset value of each data block in the fragment file.
[0273] In the above embodiments, for uploading large amounts of data, a chunked upload method can be used, which can effectively improve data upload efficiency and reduce upload time. Furthermore, users can selectively upload segments of interest, increasing upload flexibility.
[0274] In one embodiment, the device further includes:
[0275] The calling module is used to call the object supplementation or object truncation interface of the distributed file system.
[0276] The processing module is used to upload supplementary data for the target data file to the storage container or to truncate the target data file in the storage container based on the object supplementation interface or the object truncation interface; or, based on the object supplementation interface or the object truncation interface, to upload supplementary data for the target block data file to other storage containers or to truncate the target block data file in other storage containers.
[0277] In the above embodiments, the object supplementation interface allows for data supplementation of the target data file or target data file individually. Data supplementation can be performed directly within the distributed file system, eliminating the need to download and re-upload the complete target data file or target data file, thus reducing the complexity of data supplementation and improving data processing efficiency. Furthermore, the object truncation interface allows for truncation of the target data file or target data file individually. Again, data truncation can be performed directly within the distributed file system, eliminating the need to download and re-upload the complete target data file or target data file, further reducing the complexity of data supplementation and improving data processing efficiency.
[0278] In one embodiment, the device further includes:
[0279] The search module is used to find target files that meet the state transition conditions based on the object directory tree; the object directory tree is a directory tree composed of object directories and file directories, the file directories include the first file directory or the second file directory, and the target files include target data files;
[0280] The tagging module is used to mark target files as intermediate states that match state transition conditions;
[0281] The acquisition module is used to obtain a list of data for each data block in the target file after the marking is completed;
[0282] The cleanup module is used to process data blocks in the data list;
[0283] The migration module is used to migrate the target file from the intermediate state to the target state after data processing is completed.
[0284] The aforementioned data state migration method involves the server searching for target files that meet the state migration conditions based on the object directory tree. The target file is marked as an intermediate state that matches the state migration conditions. After marking, a list of data objects for each data object in the target file is obtained, and data processing is performed on the data objects in the list. This ensures that the data objects contained in the target file are processed synchronously. After all data objects have been processed, the target file is migrated from the intermediate state to the target state. This achieves synchronous data state migration of the data objects in the target file and ensures the correctness of the target file after the data state migration.
[0285] In one embodiment, the device further includes:
[0286] The receiving module is used to receive deletion requests for the target file directory under the object directory;
[0287] The calling module is used to respond to deletion requests by calling the directory deletion interface of the distributed file system;
[0288] The delete module is used to delete target file directories under the object directory based on the directory deletion interface.
[0289] The above deletion operation meets the requirement of strong data consistency. That is, after performing the deletion operation on the target data file or target block data file, the target data file or target block data file can be deleted immediately, and there is no need to search for the directories to be deleted one by one and then delete them one by one, which improves the efficiency of directory deletion.
[0290] Each module in the aforementioned data object storage device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0291] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 14 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data objects. The I / O interfaces allow the processor to exchange information with external devices. The communication interface allows communication with external terminals via a network connection. When executed by the processor, the computer program implements a data object storage method.
[0292] Those skilled in the art will understand that Figure 14 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0293] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the data object storage method described above.
[0294] In one embodiment, a computer-readable storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the steps of the data object storage method described above.
[0295] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the data object storage method described above.
[0296] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0297] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0298] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A data object storage method, characterized in that, The method includes: Create a temporary directory and an object directory in the root directory of the file system, and create a metadata-level temporary file in the temporary directory; the temporary file is a file information that is temporarily stored. The target data blocks are uploaded concurrently to the storage container to obtain the index information of the target data; When all data blocks of the target data are uploaded to the storage container concurrently, the index information of the target data is associated with the temporary file to obtain a temporary data file; when at least one data block of the target data encounters an error during the upload process, the association between the index information of the target data and the temporary file is stopped. A first file directory is created under the object directory based on the file path corresponding to the target data; The temporary data file is moved to the first file directory to obtain the target data file in the first file directory.
2. The method according to claim 1, characterized in that, The temporary directory is configured to be invisible to the user object that uploaded the target data, while the object directory is configured to be visible to the user object that uploaded the target data.
3. The method according to claim 1, characterized in that, The step of concurrently uploading data blocks of the target data to the storage container and obtaining the index information of the target data includes: The target data is divided into blocks to obtain data chunks; Each of the data blocks is uploaded concurrently to the storage container, and the block offset value corresponding to each of the data blocks is determined during the upload process. The index information of the target data is determined based on the block offset value and the block size.
4. The method according to claim 1, characterized in that, The step of associating the index information of the target data with the temporary file includes: When all data blocks of the target data are uploaded to the storage container concurrently, the index information of the target data is associated with the temporary file; The method further includes: when at least one data block of the target data encounters an anomaly during the upload process, stopping the association of the index information of the target data with the temporary file; Clean up the temporary files created in the temporary directory.
5. The method according to claim 1, characterized in that, Moving the temporary data file to the first file directory to obtain the target data file in the first file directory includes: When the target data file does not exist in the first file directory, the temporary data file is renamed to obtain the target data file; Move the target data file to the first file directory to obtain the target data file in the first file directory.
6. The method according to claim 5, characterized in that, The method further includes: When a data file with a version lower than the target version corresponding to the temporary data file exists in the first file directory, the data file in the first file directory is updated based on the temporary data file to obtain the target data file in the first file directory; If a data file with a higher version than the target data file exists in the first file directory, the move of the target data file to the first file directory is refused. The target version is the version corresponding to the temporary data file.
7. The method according to claim 1, characterized in that, After moving the temporary data file to the first file directory, the method further includes: Receive a read request for the target data file; In response to a read request for the target data file, the block size and block offset of the data blocks in the target data file are read. Based on the block size and block offset of the data block, determine the index information of the target data file; Based on the index information of the target data file, the target data file in the storage container is read.
8. The method according to claim 1, characterized in that, The method further includes: Create a metadata-level temporary block file in the temporary directory; the temporary block file is a temporary multi-segment upload file; An upload identifier is generated, and the upload identifier and the file path of the media data are associated with the temporary block file to obtain an associated file; the upload identifier is an identifier representing the temporary block file; When the media data is uploaded concurrently to other storage containers in the form of fragment files, each fragment file is associated with the associated file to obtain a temporary block data file; the fragment file is a secondary file obtained by dividing the media data into at least two segments, and the data block is a dataset of small blocks in the target data; Based on the file path corresponding to the temporary block data file, create a second file directory under the object directory; The temporary block data file is moved to the second file directory to obtain the target block data file in the second file directory.
9. The method according to claim 8, characterized in that, The method further includes: When the media data is uploaded concurrently to other storage containers in the form of fragment files, a first file offset value for each fragment file is determined; Add a fragment directory in the root directory of the file system to store the fragment files; Add a new fragment file table to store the offset values of the first file.
10. The method according to claim 9, characterized in that, Moving the temporary block data file to the second file directory includes: Move the temporary block data file corresponding to the target fragment file to the second file directory, wherein the target fragment file belongs to at least one file among the fragment files; The method further includes: determining the file offset value corresponding to the target block data file based on the first file offset value of the target fragment file.
11. The method according to claim 8, characterized in that, After moving the temporary block data file to the second file directory, the method further includes: Receive a read request for the target block data file; In response to the read request of the target block data file, a first file offset value is read from the fragment file table, and a second file offset value corresponding to the fragment file is determined; The index information of the target block data file is determined based on the first file offset value and the second file offset value; Based on the index information of the target block data file, read the target block data file from the other storage containers.
12. The method according to claim 11, characterized in that, The step of determining the second file shift value corresponding to the fragment file includes: Obtain the block size and block offset value of each data block in the fragment file; Based on the block size and block offset of each data block in the fragment file, the second file offset value corresponding to the fragment file is determined.
13. The method according to any one of claims 8 to 12, characterized in that, The method further includes: Call the object supplementation or object truncation interface of the distributed file system; Based on the object augmentation interface or the object truncation interface, augmentation data for the target data file is uploaded to the storage container, or the target data file is truncated in the storage container; or... Based on the object supplementation interface or the object truncation interface, supplementary data for the target block data file is uploaded to other storage containers, or the target block data file is truncated in other storage containers.
14. The method according to any one of claims 8 to 12, characterized in that, The method further includes: The target file that meets the state transition conditions is found based on the object directory tree; the object directory tree is a directory tree composed of the object directory and the file directory, the file directory includes the first file directory or the second file directory, and the target file includes the target data file; The target file is marked as an intermediate state that matches the state transition conditions; After marking is completed, obtain a data list for each data block in the target file; Perform data processing on the data blocks in the data list; After data processing is complete, the target file is migrated from the intermediate state to the target state.
15. The method according to any one of claims 1 to 12, characterized in that, The method further includes: Receive a deletion request for the target file directory under the object directory; In response to the deletion request, the directory deletion interface of the distributed file system is invoked; Based on the directory deletion interface, the target file directory under the object directory is deleted.
16. A data object storage device, characterized in that, The device includes: A new module has been added to create temporary directories and object directories in the root directory of the file system. The first creation module is used to create metadata-level temporary files in the temporary directory; the temporary files are temporarily stored file information. The upload module is used to concurrently upload data blocks of the target data to the storage container and obtain the index information of the target data; The association module is used to associate the index information of the target data with the temporary file to obtain a temporary data file when all data blocks of the target data are concurrently uploaded to the storage container; and to stop associating the index information of the target data with the temporary file when at least one data block of the target data encounters an error during the upload process. The second creation module is used to create a first file directory under the object directory based on the file path corresponding to the target data; The moving module is used to move the temporary data file to the first file directory to obtain the target data file in the first file directory.
17. The apparatus according to claim 16, characterized in that, The temporary directory is configured to be invisible to the user object that uploaded the target data, while the object directory is configured to be visible to the user object that uploaded the target data.
18. The apparatus according to claim 16, characterized in that, The upload module is further configured to slice the target data into blocks to obtain data blocks; upload each data block concurrently to the storage container; and determine the block offset value corresponding to each data block during the upload process. The index information of the target data is determined based on the block offset value and the block size.
19. The apparatus according to claim 16, characterized in that, The device further includes: The association module is also used to associate the index information of the target data with the temporary file when all data blocks of the target data are concurrently uploaded to the storage container; The stop module is used to stop associating the index information of the target data with the temporary file when at least one data block of the target data encounters an error during the upload process. The cleanup module is used to clean up temporary files created in the temporary directory.
20. The apparatus according to claim 16, characterized in that, The moving module is further configured to, when the target data file does not exist in the first file directory, rename the temporary data file to obtain the target data file; and move the target data file to the first file directory to obtain the target data file in the first file directory.
21. The apparatus according to claim 20, characterized in that, The device further includes: The update module is used to update the data files in the first file directory based on the temporary data file when a data file with a lower version than the target version corresponding to the temporary data file exists in the first file directory, so as to obtain the target data file in the first file directory. The stop module is used to refuse to move the target data file to the first file directory when a data file with a higher version than the target data file exists in the first file directory; The target version is the version corresponding to the temporary data file.
22. The apparatus according to claim 16, characterized in that, The device further includes: The receiving module is used to receive read requests for the target data file; The reading module is used to read the block size and block offset value of the data blocks in the target data file in response to the read request of the target data file; The determination module is used to determine the index information of the target data file based on the block size and block offset value of the data block; and to read the target data file in the storage container based on the index information of the target data file.
23. The apparatus according to claim 16, characterized in that, The device further includes: A creation module is used to create metadata-level temporary chunk files in the temporary directory; the temporary chunk files are temporary multi-segment upload files; A generation module is used to generate an upload identifier and associate the upload identifier and the file path of the media data with the temporary block file to obtain an associated file; the upload identifier is an identifier representing the temporary block file; The association module is also used to associate each of the fragment files with the association file when the media data is uploaded concurrently to other storage containers in the form of fragment files, to obtain a temporary block data file; the fragment file is a secondary file obtained by dividing the media data into at least two segments, and the data block is a dataset of small blocks in the target data; The second creation module is further configured to create a second file directory under the object directory based on the file path corresponding to the temporary block data file; The moving module is further configured to move the temporary block data file to the second file directory to obtain the target block data file in the second file directory.
24. The apparatus according to claim 23, characterized in that, The device further includes: The determination module is used to determine the first file offset value of each of the media fragment files when the media data is uploaded concurrently to other storage containers in the form of fragment files; The newly added module is also used to add a segment directory in the root directory of the file system for storing each of the segment files; and to add a segment file table for storing the offset values of the first file.
25. The apparatus according to claim 24, characterized in that, The moving module is further configured to move the temporary block data file corresponding to the target fragment file to the second file directory, wherein the target fragment file belongs to at least one of the fragment files; The determining module is further configured to determine the file offset value corresponding to the target block data file based on the first file offset value of the target fragment file.
26. The apparatus according to claim 23, characterized in that, The device further includes: The receiving module is used to receive read requests for the target block data file; The reading module is used to read a first file offset value from the fragment file table and determine a second file offset value corresponding to the fragment file in response to a read request for the target block data file; The determining module is used to determine the index information of the target block data file based on the first file offset value and the second file offset value; The reading module is also used to read the target block data file in the other storage container based on the index information of the target block data file.
27. The apparatus according to claim 26, characterized in that, The determining module is further configured to obtain the block size and block offset value of each data block in the fragment file; and determine the second file offset value corresponding to the fragment file based on the block size and block offset value of each data block in the fragment file.
28. The apparatus according to any one of claims 23 to 27, characterized in that, The device further includes: The calling module is used to call the object supplementation or object truncation interface of the distributed file system. The processing module is configured to, based on the object supplementation interface or the object truncation interface, upload supplementary data for the target data file to the storage container, or truncate the target data file in the storage container; or, based on the object supplementation interface or the object truncation interface, upload supplementary data for the target block data file to other storage containers, or truncate the target block data file in other storage containers.
29. The apparatus according to any one of claims 23 to 27, characterized in that, The device further includes: The search module is used to search for target files that meet the state transition conditions based on the object directory tree; the object directory tree is a directory tree composed of the object directory and the file directory, the file directory includes the first file directory or the second file directory, and the target file includes the target data file; A tagging module is used to tag the target file as an intermediate state that matches the state transition conditions; The acquisition module is used to acquire a data list for each data block in the target file after the marking is completed; The cleaning module is used to process the data blocks in the data list; The migration module is used to migrate the target file from the intermediate state to the target state after data processing is completed.
30. The apparatus according to any one of claims 16 to 27, characterized in that, The device further includes: A receiving module is used to receive a deletion request for the target file directory under the object directory; The calling module is used to call the directory deletion interface of the distributed file system in response to the deletion request; The deletion module is used to delete the target file directory under the object directory based on the directory deletion interface.
31. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 15.
32. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 15.
33. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 15.
Citation Information
Patent Citations
Data storage method, system and device
CN111078653A