Multi-level cache and file batch downloading method and system and storage medium

By adopting multi-level caching and file batch download methods in cloud storage systems, the problems of cloud storage systems in terms of security, availability and access speed are solved, and a more efficient, reliable and secure storage experience is achieved.

CN120201089AActive Publication Date: 2025-06-24SICHUAN LEWEI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510680156.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing cloud storage systems have problems with security, availability, convenience, and access speed, resulting in users being inefficient when using specific cloud objects to read and store.

Method used

Multi-level caching and file batch download methods are adopted to load file lists through the object storage gateway kernel, set file upload thresholds for grouping or sharding processing, and use multi-threading for file download and cache.

Benefits of technology

It improves the reliability and security of the storage system, reduces the risk of data leakage, and significantly improves the speed of read and write experience, reducing the concurrent pressure on cloud storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201089A_ABST
    Figure CN120201089A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level cache and file batch downloading method and system and a storage medium, and belongs to the technical field of computer storage. The method comprises the following steps: S1, loading a file list by an object storage gateway kernel, and assembling and de-duplicating the file list; s2, setting a file uploading threshold value, performing grouping processing on small files, and performing fragmentation processing on large files; s3, calculating the thread count required to be opened by the cloud storage gateway by using the small file grouping thread count and the large file fragmentation thread count; starting multiple threads in batches for file downloading; and S4, after obtaining the file, the cloud storage server performs caching according to the first heat threshold. Compared with the prior art, the method has the advantages that the reliability is higher, the NDN + IP-based C / S system structure is adopted, and even if the central server breaks down, the client service system is not affected when the edge network connection is normal. And the security is higher, the file is uploaded by adopting block de-duplication encryption, and compared with the traditional object storage, the risk of data leakage is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer storage technology, and in particular to a multi-level cache and file batch downloading method, system and storage medium. Background Art

[0002] The rapid development of cloud computing has led to more and more users storing data in the cloud and reading data from the cloud. However, due to the current problems of cloud storage in terms of security, availability, convenience, access speed, etc., traditional applications cannot access cloud storage well. Most public cloud providers rely on the Internet Protocol (IP), usually the REST API method on HTTP. Although it is very practical for creating new application programs, it has a very high cost for compatibility and adaptation of the original system. Due to the incompatibility between public cloud technology and original applications, it is necessary to build an object storage gateway between the cloud storage system and enterprise applications. The object storage gateway needs to have the characteristics of local high-speed caching, ensuring data security, and easy use. The price of cloud storage is more advantageous than that of traditional storage, and it can help enterprises easily go to the cloud through the object storage gateway. The object storage gateway provides basic protocol conversion and simple connectivity, so that incompatible technologies can communicate transparently. The gateway can make object storage appear like a NAS filter, block storage array, local disk, backup target, or even an extension of the application itself.

[0003] Currently, users have low efficiency when using specific cloud objects for reading and storage. The main reason for the problem is that cloud storage uses centralized servers, and related needs can only be met after interacting with the centralized server in the network center. Specifically, first, the pipeline is single-source and single-path, which is easy to cause congestion. For example, the same video needs to be sent from a single central server countless times. Second, when the business system and the object storage network fail, the normal use of the business system is seriously affected. Therefore, the present invention proposes a communication method for efficient storage, which makes minimal changes to the existing system to solve the above problems. Summary of the invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and to provide a multi-level cache and file batch downloading method, system and storage medium.

[0005] The object of the present invention is achieved through the following technical solutions: The first aspect of the present invention provides: a multi-level cache and file batch downloading method, comprising the following steps: S1: The object storage gateway kernel loads the file list, assembles and deduplicates the file list; S2: Set the file upload threshold. Files smaller than the file upload threshold are small files, and those larger than the file upload threshold are large files. Group the small files and fragment the large files. S3: Calculate the number of threads that the cloud storage gateway needs to start by adding the number of small file grouping threads and the number of large file fragmenting threads. Batch start multiple threads for file download. S4: After the cloud storage server obtains the file, cache it according to the first heat threshold.

[0006] Preferably, the S1 further includes the following steps: The object storage gateway kernel loads the file list, mounts the local Posix disk through the object storage gateway, and schedules with the operating system kernel to obtain the file list, including the file corresponding hash value, file size, file attributes, and initiate file download. Perform file list assembly and deduplication, calculate the total number of files, perform integrated analysis based on the file hash value, analyze the file duplication situation, and determine whether there are duplicate files in the files. If there are duplicate file hash values, remove the duplicate files from the download list and record the duplicate file information locally.

[0007] Preferably, the S2 further includes the following steps: First, perform small file upload grouping according to the number of files and the file upload threshold. By looping through the file list, add up the file sizes to initially calculate the grouping size. When the target download file is a large file, calculate the download fragmentation rule. The number of download fragments = the size of the target download file / the file download threshold. Add 1 to the number of groups and recalculate the grouping size. When the target upload file is a large file, calculate the upload fragmentation rule. The number of upload fragments = the sum of the sizes of the target upload files / the file upload threshold. When the size of this batch of files to be uploaded is greater than the file upload threshold, add 1 to the number of groups. Calculate the overall grouping size through looping.

[0008] Preferably, the S3 further includes the following steps: After calculating the number of threads that the cloud storage gateway needs to start, batch start multiple threads for file download according to the number of threads that the storage gateway client can start. The upload thread batch checks whether there are files with the corresponding hash value in the local disk cache according to the file hash value. If it exists in the local disk cache, directly use the local io to open the file, and the cloud storage gateway returns the file stream to the application through the system kernel. If it does not exist in the local disk cache, the storage gateway initiates a file batch download request to the cloud storage gateway server. When the cloud storage server receives a batch download request, it circularly checks the download file list; compares the hash values of the files in the download file list with the files in the redis cache; if the cloud storage server has the file, it reads the file stream in the high-speed temporary buffer for output, records the file download times and download times, and writes them into the redis cache; if the cloud storage server does not have the file, it sends a file download request to the object storage gateway; the object storage gateway, according to its own business logic, obtains the file storage location through the crash algorithm, and returns the file stream through ssd cache acceleration for the cloud storage server to download.

[0009] Preferably, step S4 further includes the following steps: After the cloud storage server obtains the file, it determines whether to cache according to the file information recorded locally; if the heat value is higher than the first heat threshold, the file is downloaded to the storage gateway cache area; if the heat value is lower than the first heat threshold, it is directly input to the storage gateway client through the proxy method; The storage gateway server combines all the file streams into a packaged file stream according to the file size range information of the returned file streams; records the total data size of the packaged file stream, the size of each file, and the corresponding order of the files in the returned file metadata information; The storage gateway client obtains the packaged file stream, parses it and downloads it to the local disk. The local disk starts multiple threads according to the number of files, and splits the stream file according to the file metadata size, file order, and pointer position; Each split file is restored according to the corresponding metadata information and sent to the application stream through kernel scheduling; The local disk records the corresponding hash value of the file, the file download times, and the file download time; The local cache detection process calculates the file heat value according to the file download times, file call times, and file download time; According to the set local cache space size and the file heat value sorting of each file, asynchronously deletes the cache files with file heat values lower than the second heat threshold.

[0010] The second aspect of the present invention provides: A multi-level cache and file batch download system for implementing any of the above multi-level cache and file batch download methods, including: A loading module for using the object storage gateway kernel to load the file list and assembling and deduplicating the file list; A processing module for setting a file upload threshold, files smaller than the file upload threshold are small files, and files larger than the file upload threshold are large files, grouping small files and fragmenting large files; A download module, which is used to calculate the number of threads that need to be enabled by the cloud storage gateway by adding the number of small file grouping threads and the number of large file sharding threads; batch enable multiple threads to download files. A caching module, which is used to cache according to the first heat threshold after obtaining a file from the cloud storage server.

[0011] The third aspect of the present invention provides: a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, any of the above multi-level caching and file batch download methods is implemented.

[0012] The beneficial effects of the present invention are: 1) Stronger reliability. The system structure of NDN + IP-based C / S is adopted. Even if the central server fails, the customer business system is not affected at all when the edge network connection is normal.

[0013] 2) Stronger security. The file is uploaded in a deduplicated and encrypted manner in blocks. Compared with traditional object storage, the risk of data leakage is greatly reduced.

[0014] 3) Higher read and write experience speed. Based on the design of local hot and cold cache data, the read and write speed depends on the local cache device. At the same time, the concurrent pressure on the Tianyi Cloud object storage is also reduced. Because of the cache design of the object storage gateway, the interaction pressure between the customer and the object storage is reduced to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of the multi-level caching and file batch download method; Figure 2 It is a detailed flowchart of the multi-level caching and file batch download method; Figure 3 It is a flowchart of large file sharding download; Figure 4 It is a flowchart of small file grouping download; Figure 5 It is a flowchart of cache file cleaning calculation; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0017] The present invention adopts the NDN (Name Data Network) + Posix-based C / S (Client / Server) architecture to implement the conversion between the two protocols through this method. This method is based on the Posix-based C / S architecture and improves the response speed of the cloud storage service through mechanisms such as local data packet fragmentation, multi-level caching, and concurrent uploading. This method utilizes the integration of cloud storage and the network to establish a low-cost, scalable, secure, and convenient integrated cloud storage platform that combines local high-speed caching and cloud storage.

[0018] Refer to Figures 1 - 5 , the first aspect of the present invention provides: a multi-level caching and file batch downloading method, including the following steps: S1: The object storage gateway kernel loads the file list and assembles and deduplicates the file list; S2: Set the file upload threshold. Files smaller than the file upload threshold are small files, and files larger than the file upload threshold are large files. Group the small files and fragment the large files; S3: Calculate the number of threads required to be started by the cloud storage gateway by adding the number of small file grouping threads and the number of large file fragmenting threads; Batch start multiple threads to download files; S4: After the cloud storage server obtains the file, cache it according to the first heat threshold.

[0019] In this embodiment, the client adopts an IP-based C / S architecture, improves the response speed of the cloud storage service through the local data caching mechanism, and realizes local network high-speed storage through hardware and software configuration. Three-level caching of gateway local storage, gateway server storage, and object storage components is used to achieve multi-level caching acceleration for downloading. Each level of cache calculates the heat value according to the file download time, file download times, file call times, etc. The object storage gateway adopts the NDN (Name Data Network) + IP-based C / S (Client / Server) architecture to implement the conversion between the two protocols at the gateway. Using NDN network storage as the backend, the object storage gateway uploads data to the NDN network storage in an encrypted manner.

[0020] In some embodiments, the S1 further includes the following steps: The object storage gateway kernel loads the file list, mounts the local Posix disk through the object storage gateway, schedules with the operating system kernel by the object storage gateway to obtain the file list, including the file corresponding hash value, file size, file attributes, and initiate file download; Assemble and deduplicate the file list, calculate the total number of files, perform integrated analysis based on the file hash values, analyze the file duplication situation, and determine whether there are duplicate files in the files; if there are duplicate file hash values, remove the duplicate files from the download list and record the duplicate file information locally.

[0021] In some embodiments, step S2 further includes the following steps: First, group the small file uploads according to the number of files and the file upload threshold. By looping through the file list, add up the file sizes to initially calculate the group size. When the target download file is a large file, calculate the download sharding rule. The number of download shards = the size of the target download file / the file download threshold. Increment the number of groups by 1 and recalculate the group size. When the target upload file is a large file, calculate the upload sharding rule. The number of upload shards = the sum of the sizes of the target upload files / the file upload threshold. When the upload size of this batch of files is greater than the file upload threshold, increment the number of groups by 1. Calculate the overall group size through looping.

[0022] In some embodiments, step S3 further includes the following steps: After calculating the number of threads required to be enabled by the cloud storage gateway, batch enable multiple threads for file download according to the number of threads that can be enabled by the storage gateway client. The upload threads batch detect whether there are files with corresponding hash values in the local disk cache according to the file hash values; if the local disk cache exists, directly use the local io to open the file, and the cloud storage gateway returns the file stream to the application through the system kernel; if it does not exist in the local disk cache, the storage gateway sends a batch file download request to the cloud storage gateway server. When the cloud storage server receives the batch download request, loop through and detect the download file list; compare the hash values of the files in the download file list with the files in the redis cache; if the file exists in the cloud storage server, read the file stream from the high-speed temporary buffer for output, record the file download times and download time, and write them into the redis cache; if the file does not exist in the cloud storage server, send a file download request to the object storage gateway; the object storage gateway, according to its own business logic, obtains the file storage location through the crash algorithm, accelerates through the ssd cache, and returns the file stream for the cloud storage server to download.

[0023] In some embodiments, step S4 further includes the following steps: After the cloud storage server obtains a file, it determines whether caching is required based on the file information recorded locally; if the popularity value is higher than the first popularity threshold, the file is downloaded to the storage gateway cache area; if the popularity value is lower than the first popularity threshold, it is directly input into the storage gateway client through the proxy method; Based on the returned file streams, the storage gateway server combines all the file streams into a packaged file stream according to the file size range information; the returned file metadata information records the total data size of the packaged file stream, the size of each file, and the corresponding order of the files; The storage gateway client obtains the packaged file stream, parses it and downloads it to the local disk. The local disk starts multiple threads according to the number of files, and splits the stream file according to the file metadata size, file order, and pointer position; Each split file is restored according to the corresponding metadata information and sent to the application stream through kernel scheduling; The local disk records the corresponding hash value of the file, as well as the file download times and file download time; The local cache detection process calculates the file popularity value based on the file download times, file call times, and file download time; Sorted according to the set local cache space size and the file popularity value of each file, cache files with a file popularity value lower than the second popularity threshold are asynchronously deleted.

[0024] Two embodiments are provided below: Example 1: Video on demand.

[0025] The client uses an IP connection to upload a data packet containing video on demand request information (such as the name of the video on demand file). The target IP address is a pre-agreed server address. The object storage gateway receives the data packet downloaded by the client, removes the IP-related packet header at the application layer and parses it to obtain the video on demand request information. The gateway looks up the relevant block numbers according to the video on demand file name, generates the corresponding hash check code, and checks the local routing table. If the corresponding data is local, it directly uses the IP connection to send the data to the client. If the corresponding data is in other storage gateways, an interest packet (containing the required data information) in the NDN protocol is sent to the other storage gateways, and the other storage gateways reply with the corresponding data packet. The storage gateway that receives the data then uses the IP connection to send the on-demand data to the client.

[0026] Example 2: The client reads its own stored data.

[0027] The client calculates the hash value of the data for obtaining the file list to remove duplicates. The user can customize the block size for concurrent downloads, and each file is verified using a hash. If the file block exists in the local cache, the file block does not need to be downloaded again, and only the download count of the file block is recorded, reducing the consumption of network I / O by the file. The storage gateway receives the data packet requested by the client, parses it after removing the IP-related packet header at the application layer, and performs relevant proxy forwarding according to the client request. If it is necessary to download a certain data / data shard, the object storage gateway records the data name and shard number, generates the corresponding hash verification code based on the data name and shard number, and searches in this gateway or the adjacent gateway. The storage gateway client synchronously establishes a routing table, calculates the file popularity value based on the download time, file size, download count, etc. of the file, and the storage gateway associates the file path with the corresponding hash to establish a popularity value routing table.

[0028] The second aspect of the present invention provides: A multi-level cache and file batch download system for implementing any of the above multi-level cache and file batch download methods, including: A loading module for loading the file list using the object storage gateway kernel and assembling and removing duplicates from the file list; A processing module for setting a file upload threshold. Files with a size smaller than the file upload threshold are small files, and those larger than the file upload threshold are large files. Group the small files and fragment the large files; A download module for calculating the number of threads that the cloud storage gateway needs to start by adding the number of small file grouping threads and the number of large file fragmenting threads; batch start multiple threads to download files; A caching module for caching according to the first popularity threshold after obtaining the file on the cloud storage server side.

[0029] The third aspect of the present invention provides: A computer-readable storage medium in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, any of the above multi-level cache and file batch download methods are implemented.

[0030] The above is only the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, and can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in the relevant field. And the changes and modifications made by those skilled in the art that do not depart from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.

Claims

1. A multi-level cache and file batch download method, characterized in that: It includes the following steps: S1: The object storage gateway kernel loads the file list, and assembles and deduplicates the file list; S2: Set the file upload threshold. Files smaller than the file upload threshold are small files, and files larger than the file upload threshold are large files. Group the small files and fragment the large files; S3: Calculate the number of threads that the cloud storage gateway needs to start using the number of small file grouping threads plus the number of large file fragmentation threads; Batch start multiple threads for file download; S4: After the cloud storage server obtains the file, cache it according to the first heat threshold.

2. The multi-level cache and file batch download method according to claim 1, characterized in that: The S1 also includes the following steps: The object storage gateway kernel loads the file list, mounts the local Posix disk through the object storage gateway, schedules between the object storage gateway and the operating system kernel, obtains the file list, including the file corresponding hash value, file size, file attributes and initiates a file download; Perform file list assembly and deduplication, calculate the total number of files, perform integrated analysis according to the file hash value, analyze the file duplication situation, and judge whether there are duplicate files in the files; If there are duplicate file hash values, remove the duplicate files from the download list and record the duplicate file information locally.

3. The multi-level cache and file batch download method according to claim 1, wherein: The S2 also includes the following steps: First, group the small file uploads according to the number of files and the file upload threshold. By looping through the file list, add up the file sizes and initially calculate the group size; When the target download file is a large file, calculate the download fragmentation rule. The number of download fragments = the size of the target download file / the file download threshold, add 1 to the number of groups, and recalculate the group size at the same time; When the target upload file is a large file, calculate the upload fragmentation rule. The number of upload fragments = the sum of the sizes of the target upload files / the file upload threshold. When the size of this batch of file uploads is greater than the file upload threshold, add 1 to the number of groups; Calculate the overall group size through looping.

4. The multi-level cache and file batch download method according to claim 1, wherein: The S3 also includes the following steps: After calculating the number of threads that the cloud storage gateway needs to start, batch start multiple threads for file download according to the number of threads that the storage gateway client can start; The upload thread batch checks whether there are files with the corresponding hash value in the local disk cache according to the file hash value; If it exists in the local disk cache, directly use the local io to open the file, and the cloud storage gateway returns the file stream to the application through the system kernel; If it does not exist in the local disk cache, the storage gateway initiates a file batch download request to the cloud storage gateway server; When the cloud storage server receives a batch download request, it cyclically checks the download file list; compares the hash values of the files in the download file list with the files in the redis cache; if the file exists in the cloud storage server, it reads the file stream in the high-speed temporary buffer for output, records the file download times and download time, and writes them into the redis cache; if the file does not exist in the cloud storage server, it sends a file download request to the object storage gateway; the object storage gateway, according to its own business logic, obtains the file storage location through the crash algorithm, and returns the file stream through ssd cache acceleration for the cloud storage server to download.

5. The multi-level cache and file batch download method according to claim 1, characterized in that: Step S4 further includes the following steps: After the cloud storage server obtains the file, it determines whether caching is required according to the file information recorded locally; if the heat value is higher than the first heat threshold, it downloads the file to the storage gateway cache area; if the heat value is lower than the first heat threshold, it directly inputs the storage gateway client through the proxy method; The storage gateway server combines all the file streams into a packaged file stream according to the file size range information of the file streams returned; records the total data size of the packaged file stream, the size of each file, and the corresponding order of the files in the returned file metadata information; The storage gateway client obtains the packaged file stream, parses it and downloads it to the local disk. The local disk starts multiple threads according to the number of files, and splits the stream file according to the file metadata size, file order, and pointer position; Restores each split file according to the corresponding metadata information, and schedules it to the application stream through the kernel; The local disk records the corresponding hash value of the file, as well as the file download times and file download time; The local cache detection process calculates the file heat value according to the file download times, file call times, and file download time; According to the set local cache space size and the file heat value sorting of each file, asynchronously deletes the cached files with file heat values lower than the second heat threshold.

6. A multi-level cache and file batch download system, characterized in that: For implementing the multi-level cache and file batch download method according to any one of claims 1-5, including: A loading module, used to load the file list using the object storage gateway kernel, and assemble and deduplicate the file list; A processing module, used to set the file upload threshold, files with a size smaller than the file upload threshold are small files, and files larger than the file upload threshold are large files, group the small files, and slice the large files; A download module, used to calculate the number of threads that the cloud storage gateway needs to start by adding the number of small file grouping threads and the number of large file slicing threads; batch start multiple threads for file download; A caching module, used to cache according to the first heat threshold after the cloud storage server obtains the file.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the multi-level cache and file batch download method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Large file block uploading and encrypted storage method based on fastdfs

    CN116389461A

  • File uploading method and device, equipment and storage medium

    CN117319375A

  • System with distributed storage structure

    CN118018563A

  • File uploading and downloading method and system

    CN118748672A

  • Cache-based file uploading cost optimization method, device and equipment

    CN118916334A